Skip to content

v0.14

Choose a tag to compare

@jararap jararap released this 17 Apr 04:12
· 7 commits to main since this release

Major changes (from 0.13):

  • Implement greedy updates, speeding up the algorithm by 20x-100x
  • Split greedy_builder into greedy_encoder and pco_tokenizer
  • Add huggingface transformer friendly module with similar API (pcatt.hf)
  • Change some class methods to return py::bytes instead of string to avoid invalid UTF-8

From beta8 to 0.14

  • Add verbosity option
  • Allow default padding and truncating behaviours from huggingface PreTrainedTokenizer
  • fix tokenize and vocab_size calls

From beta7 to beta8:

  • fix another bug where there is id mismatch amongst special tokens
  • align tokenizer behaviour to HF API: some properties declared in constructor becomes the defaults in call

From beta6 to beta7:

  • Fix a bug where there are id mismatch for special tokens. May have affected pretraining.

From beta5 to beta6:

  • optimize alter_graph
  • fix unknown bug (whose impact is extremely minor)

From beta4 to beta5:

  • add decode/batch_decode to greedy_encoder
  • add parallelization for token candidate generation for pco_tokenizer (rolled back)
  • building the token set is now twice as fast as before

From beta3 to beta4:

  • Fix encode to take in str input
  • Add ability to pass callbacks

From beta to beta3:

  • Fix behaviour of train_new_from_iterator to take list[str], while maintaining ability to process list[list[str]]
  • Additionally, allow users to define their desired regex pattern to use in regex.findall