Skip to content

sinlib v0.2.0

Latest

Choose a tag to compare

@ransaka ransaka released this 23 Jul 11:46
· 14 commits to main since this release

1. Hybrid Spelling Correction Engine

  • Phonological Trie Search (Retrieval Fallback): A prefix tree built over Sinhala phonological tokens (Aksharas). Computes edit distance on token sequences with custom keyboard-adjacency and phonetic confusion weights (e.g. ප $\rightarrow$ පා cost 0.4).
  • Beam Search Seq2Seq Corrector (Generative): A Bidirectional GRU Encoder + GRU Decoder with Attention, utilizing Beam Search Decoding (width=3) to prevent repeating character loops.
  • Stupid Backoff Language Model: Utilizes bigram and unigram probabilities derived from a $151\text{k}$-article Sinhala news corpus. Automatically backs off to unigram priors for unseen words, maintaining correction accuracy for descriptive words (e.g. ලස්සන).
  • Vocabulary Expansion: Extended vocabulary coverage from $20\text{k}$ to $61,436$ words to cover classical verb inflections (ගියෙමි).

2. Pretrained Typo Detection

  • Statistical Akshara Trigram Model: Replaces character-level bigrams with a trigram model built over tokenized sequences, mapped mathematically to preserve the $10^{-8}$ threshold.
  • PyTorch Bi-GRU Sequence Labeler: Pinpoints the exact location of spelling errors in a word (e.g. ය් $\rightarrow$ යි).
  • Zero-Dependency Fallback: The library seamlessly runs spelling checks using the Trigram model if PyTorch (torch) is absent.

3. Tokenizer & Preprocessing Upgrades

  • Tensor Returns: Tokenizer.__call__ and Tokenizer.encode_plus now support return_tensors="pt" (PyTorch), "tf" (TensorFlow), and "np" (NumPy).
  • Subword Tokenizer: A new SubwordTokenizer class that builds subword tokens on top of Sinhala phonology.
  • Sinhala Normalizer: Added normalize_sinhala(text) to standardize ZWJ/ZWNJ layout sequences and enforce Unicode Normalization Form C (NFC).