You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Phonological Trie Search (Retrieval Fallback): A prefix tree built over Sinhala phonological tokens (Aksharas). Computes edit distance on token sequences with custom keyboard-adjacency and phonetic confusion weights (e.g. ප$\rightarrow$පා cost 0.4).
Beam Search Seq2Seq Corrector (Generative): A Bidirectional GRU Encoder + GRU Decoder with Attention, utilizing Beam Search Decoding (width=3) to prevent repeating character loops.
Stupid Backoff Language Model: Utilizes bigram and unigram probabilities derived from a $151\text{k}$-article Sinhala news corpus. Automatically backs off to unigram priors for unseen words, maintaining correction accuracy for descriptive words (e.g. ලස්සන).
Vocabulary Expansion: Extended vocabulary coverage from $20\text{k}$ to $61,436$ words to cover classical verb inflections (ගියෙමි).
2. Pretrained Typo Detection
Statistical Akshara Trigram Model: Replaces character-level bigrams with a trigram model built over tokenized sequences, mapped mathematically to preserve the $10^{-8}$ threshold.
PyTorch Bi-GRU Sequence Labeler: Pinpoints the exact location of spelling errors in a word (e.g. ය්$\rightarrow$යි).
Zero-Dependency Fallback: The library seamlessly runs spelling checks using the Trigram model if PyTorch (torch) is absent.
3. Tokenizer & Preprocessing Upgrades
Tensor Returns: Tokenizer.__call__ and Tokenizer.encode_plus now support return_tensors="pt" (PyTorch), "tf" (TensorFlow), and "np" (NumPy).
Subword Tokenizer: A new SubwordTokenizer class that builds subword tokens on top of Sinhala phonology.
Sinhala Normalizer: Added normalize_sinhala(text) to standardize ZWJ/ZWNJ layout sequences and enforce Unicode Normalization Form C (NFC).