Skip to content

Releases: ransaka/sinlib

sinlib v0.2.0

Choose a tag to compare

@ransaka ransaka released this 23 Jul 11:46

1. Hybrid Spelling Correction Engine

  • Phonological Trie Search (Retrieval Fallback): A prefix tree built over Sinhala phonological tokens (Aksharas). Computes edit distance on token sequences with custom keyboard-adjacency and phonetic confusion weights (e.g. ප $\rightarrow$ පා cost 0.4).
  • Beam Search Seq2Seq Corrector (Generative): A Bidirectional GRU Encoder + GRU Decoder with Attention, utilizing Beam Search Decoding (width=3) to prevent repeating character loops.
  • Stupid Backoff Language Model: Utilizes bigram and unigram probabilities derived from a $151\text{k}$-article Sinhala news corpus. Automatically backs off to unigram priors for unseen words, maintaining correction accuracy for descriptive words (e.g. ලස්සන).
  • Vocabulary Expansion: Extended vocabulary coverage from $20\text{k}$ to $61,436$ words to cover classical verb inflections (ගියෙමි).

2. Pretrained Typo Detection

  • Statistical Akshara Trigram Model: Replaces character-level bigrams with a trigram model built over tokenized sequences, mapped mathematically to preserve the $10^{-8}$ threshold.
  • PyTorch Bi-GRU Sequence Labeler: Pinpoints the exact location of spelling errors in a word (e.g. ය් $\rightarrow$ යි).
  • Zero-Dependency Fallback: The library seamlessly runs spelling checks using the Trigram model if PyTorch (torch) is absent.

3. Tokenizer & Preprocessing Upgrades

  • Tensor Returns: Tokenizer.__call__ and Tokenizer.encode_plus now support return_tensors="pt" (PyTorch), "tf" (TensorFlow), and "np" (NumPy).
  • Subword Tokenizer: A new SubwordTokenizer class that builds subword tokens on top of Sinhala phonology.
  • Sinhala Normalizer: Added normalize_sinhala(text) to standardize ZWJ/ZWNJ layout sequences and enforce Unicode Normalization Form C (NFC).

0.1.11

Choose a tag to compare

@ransaka ransaka released this 03 Jan 04:36
a20a5b7

Full Changelog: 0.1.10...0.1.11

0.1.10: fix(tokenizer): update token count assertion and handle unknown tokens test(tokenizer)

Choose a tag to compare

@ransaka ransaka released this 10 Aug 04:54

fix(tokenizer): update token count assertion and handle unknown tokens

  • test(tokenizer): update tests to reflect unknown token handling
  • docs: add note about temporarily unavailable modules
  • chore: update gitignore and disable romanizer tests

0.1.9.3: feat(tokenizer): add BOS token and batch processing support

Choose a tag to compare

@ransaka ransaka released this 12 Apr 09:00
Add support for a beginning-of-sequence (BOS) token in the Tokenizer class, including initialization, encoding, decoding, and batch processing. This allows for more flexible tokenization workflows, particularly for sequence-based models. Additionally, implement batch encoding and decoding methods to improve efficiency when processing multiple texts simultaneously.

The changes include:
- Adding a BOS token to the Tokenizer class.
- Extending the encode and decode methods to handle BOS tokens.
- Introducing batch_encode and batch_decode methods for processing multiple texts in a single call.
- Updating documentation and adding corresponding unit tests.

0.1.9.2: feat(tokenizer): add attention mask support and reorder special tokens

Choose a tag to compare

@ransaka ransaka released this 12 Apr 07:00
This commit introduces support for returning attention masks during tokenization, which is essential for models that require masking padded tokens. Additionally, the special tokens have been reordered to ensure consistent IDs (pad_token_id=0, unknown_token_id=1, end_of_text_token_id=2). The changes improve the tokenizer's compatibility with transformer-based models.

0.1.9.1

Choose a tag to compare

@ransaka ransaka released this 29 Mar 18:05

Full Changelog: 0.1.9...0.1.9.1

0.1.9

Choose a tag to compare

@ransaka ransaka released this 23 Mar 18:03

Full Changelog: 0.1.8...0.1.9

0.1.8

Choose a tag to compare

@ransaka ransaka released this 23 Mar 13:36

Update the version in __init__.py and pyproject.toml to 0.1.8. Remove unused data files (config.json, sinhala_chars_with_special_chars.txt, char_map.json, vocab.json) and update the logo reference in README.md. Add return statement in tokenizer.py to prevent unintended execution.

0.1.7

Choose a tag to compare

@ransaka ransaka released this 23 Mar 12:58

Full Changelog: 0.1.6...0.1.7

0.1.6

Choose a tag to compare

@ransaka ransaka released this 23 Mar 12:34

Full Changelog: 0.1.5...0.1.6