Skip to content

0.1.9.3: feat(tokenizer): add BOS token and batch processing support

Choose a tag to compare

@ransaka ransaka released this 12 Apr 09:00
· 33 commits to main since this release
Add support for a beginning-of-sequence (BOS) token in the Tokenizer class, including initialization, encoding, decoding, and batch processing. This allows for more flexible tokenization workflows, particularly for sequence-based models. Additionally, implement batch encoding and decoding methods to improve efficiency when processing multiple texts simultaneously.

The changes include:
- Adding a BOS token to the Tokenizer class.
- Extending the encode and decode methods to handle BOS tokens.
- Introducing batch_encode and batch_decode methods for processing multiple texts in a single call.
- Updating documentation and adding corresponding unit tests.