Releases
v0.14
Compare
Sorry, something went wrong.
No results found
Major changes (from 0.13):
Implement greedy updates, speeding up the algorithm by 20x-100x
Split greedy_builder into greedy_encoder and pco_tokenizer
Add huggingface transformer friendly module with similar API (pcatt.hf)
Change some class methods to return py::bytes instead of string to avoid invalid UTF-8
From beta8 to 0.14
Add verbosity option
Allow default padding and truncating behaviours from huggingface PreTrainedTokenizer
fix tokenize and vocab_size calls
From beta7 to beta8:
fix another bug where there is id mismatch amongst special tokens
align tokenizer behaviour to HF API: some properties declared in constructor becomes the defaults in call
From beta6 to beta7:
Fix a bug where there are id mismatch for special tokens. May have affected pretraining.
From beta5 to beta6:
optimize alter_graph
fix unknown bug (whose impact is extremely minor)
From beta4 to beta5:
add decode/batch_decode to greedy_encoder
add parallelization for token candidate generation for pco_tokenizer (rolled back)
building the token set is now twice as fast as before
From beta3 to beta4:
Fix encode to take in str input
Add ability to pass callbacks
From beta to beta3:
Fix behaviour of train_new_from_iterator to take list[str], while maintaining ability to process list[list[str]]
Additionally, allow users to define their desired regex pattern to use in regex.findall
You can’t perform that action at this time.