Version 0.3.0 adds selective language routing and the first measured benchmark for the package.
The routing API can reject decisions below configurable confidence or score-margin thresholds while retaining the underlying evidence. Coverage-aware evaluation reports accuracy and coverage for accepted routes, with per-language results.
The AfriSenti benchmark pins the official source revision and file hashes for five languages: Hausa, Igbo, Nigerian Pidgin, Swahili and Yoruba. A policy selected on development data reached 89.65% accuracy at 74.03% coverage on 18,402 test examples. The paired orthography test found that removing diacritics reduced overall accuracy from 74.75% to 69.87%.
Source tweets are downloaded locally and are not included in this repository or its reports.
Method and per-language results, including limitations: https://oyinkanchekwas.github.io/low-resource-nlp-toolkit/benchmark/