Skip to content

Repository files navigation

Tokenizer efficiency on traditional chinese sequence

This repository showcases the tokenizer efficiencies of several existing LLMs on Traditional Chinese sequences. Based on the definition of tokenizer efficiency in technical report of Bailong, we utilize tokenizers to tokenize the sequences in Traditional Chinese Universal Dependencies Treebank and compute the tokenizer efficiency. The detailed implementation and main results are presented in Efficiency_test.ipynb.

Citation

@misc{chen2024bailong,
      title={Bailong: Bilingual Transfer Learning based on QLoRA and Zip-tie Embedding}, 
      author={Lung-Chuan Chen and Zong-Ru Li},
      year={2024},
      eprint={2404.00862},
      archivePrefix={arXiv},
      primaryClass={cs.CL}
}

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages