Releases: ransaka/sinlib
Releases · ransaka/sinlib
Release list
sinlib v0.2.0
1. Hybrid Spelling Correction Engine
-
Phonological Trie Search (Retrieval Fallback): A prefix tree built over Sinhala phonological tokens (Aksharas). Computes edit distance on token sequences with custom keyboard-adjacency and phonetic confusion weights (e.g.
ප$\rightarrow$ පාcost0.4). - Beam Search Seq2Seq Corrector (Generative): A Bidirectional GRU Encoder + GRU Decoder with Attention, utilizing Beam Search Decoding (width=3) to prevent repeating character loops.
-
Stupid Backoff Language Model: Utilizes bigram and unigram probabilities derived from a
$151\text{k}$ -article Sinhala news corpus. Automatically backs off to unigram priors for unseen words, maintaining correction accuracy for descriptive words (e.g.ලස්සන). -
Vocabulary Expansion: Extended vocabulary coverage from
$20\text{k}$ to$61,436$ words to cover classical verb inflections (ගියෙමි).
2. Pretrained Typo Detection
-
Statistical Akshara Trigram Model: Replaces character-level bigrams with a trigram model built over tokenized sequences, mapped mathematically to preserve the
$10^{-8}$ threshold. -
PyTorch Bi-GRU Sequence Labeler: Pinpoints the exact location of spelling errors in a word (e.g.
ය්$\rightarrow$ යි). -
Zero-Dependency Fallback: The library seamlessly runs spelling checks using the Trigram model if PyTorch (
torch) is absent.
3. Tokenizer & Preprocessing Upgrades
- Tensor Returns:
Tokenizer.__call__andTokenizer.encode_plusnow supportreturn_tensors="pt"(PyTorch),"tf"(TensorFlow), and"np"(NumPy). - Subword Tokenizer: A new
SubwordTokenizerclass that builds subword tokens on top of Sinhala phonology. - Sinhala Normalizer: Added
normalize_sinhala(text)to standardize ZWJ/ZWNJ layout sequences and enforce Unicode Normalization Form C (NFC).
0.1.11
0.1.10: fix(tokenizer): update token count assertion and handle unknown tokens test(tokenizer)
fix(tokenizer): update token count assertion and handle unknown tokens
- test(tokenizer): update tests to reflect unknown token handling
- docs: add note about temporarily unavailable modules
- chore: update gitignore and disable romanizer tests
0.1.9.3: feat(tokenizer): add BOS token and batch processing support
Add support for a beginning-of-sequence (BOS) token in the Tokenizer class, including initialization, encoding, decoding, and batch processing. This allows for more flexible tokenization workflows, particularly for sequence-based models. Additionally, implement batch encoding and decoding methods to improve efficiency when processing multiple texts simultaneously. The changes include: - Adding a BOS token to the Tokenizer class. - Extending the encode and decode methods to handle BOS tokens. - Introducing batch_encode and batch_decode methods for processing multiple texts in a single call. - Updating documentation and adding corresponding unit tests.
0.1.9.2: feat(tokenizer): add attention mask support and reorder special tokens
This commit introduces support for returning attention masks during tokenization, which is essential for models that require masking padded tokens. Additionally, the special tokens have been reordered to ensure consistent IDs (pad_token_id=0, unknown_token_id=1, end_of_text_token_id=2). The changes improve the tokenizer's compatibility with transformer-based models.
0.1.9.1
0.1.9
0.1.8
Update the version in __init__.py and pyproject.toml to 0.1.8. Remove unused data files (config.json, sinhala_chars_with_special_chars.txt, char_map.json, vocab.json) and update the logo reference in README.md. Add return statement in tokenizer.py to prevent unintended execution.