Plain-text he Wikipedia lines used as the MLM pre-training corpus for rababa’s he diacritization model.
For Arabic, this augments the gold Tashkeela fine-tune corpus with ~100,000 lines of unpointed Modern Standard Arabic prose.
For Hebrew, this is the unpointed source corpus for distillation —
the rababa Modal distillation pipeline runs each line through the
Dicta Nakdan API to produce pointed labels. The distilled result lives
in a separate rababa-hebrew-distilled repo.
Fetched from the
wikimedia/wikipedia
dataset on Hugging Face (he config, 20231101 dump) via
scripts/fetch_wiki_corpus.py in the main
rababa repo.
Wikipedia text is © Wikipedia contributors, licensed under CC-BY-SA 4.0. This compiled corpus follows the same license.