Skip to content

ara-diac-small-2.0

Choose a tag to compare

@ronaldtse ronaldtse released this 30 Aug 09:55
· 155 commits to main since this release

ara-diac-small-2.0

IMF v1 (fp32, decoder kv, opset 14). Trained from sequence-level KD from the r7 canonical teacher (rababa_arabic_byt5/run-007-news/best, 2.2864 windowed DER-CE full protocol): fresh greedy r7 labels on the same r5-units corpus/limits as ara-diac-small-1.0, Muon optimizer (E3-adopted), vanilla ByT5-small (E4, pre-registered gate <= 6.26). Checkpoint rababa-checkpoints:/rababa_arabic_distill_small/run-006-r7-muon/best. The two measured wins compound: 8.259 -> 4.8218 full-set windowed DER-CE (teacher reproduces 2.289 in-run vs documented 2.2864) — a 42% error reduction on the 1.0 release at the same architecture and artifact size. Still misses the strict teacher+0.5pp gate (+2.53pp; miss disclosed); the E2/E3 factorial attributes the residual to domain coverage..

field value
task diacritization (Arab → Arab)
artifact clean-fp32.zip (1.32 GiB)
der_teacher_fullset 2.289 — windowed DER-CE (1400-byte windows, word-boundary split, greedy, haraqat-projected, Misraj evaluator); full 1,200-paragraph SadeedDiac-25; in-run reproduction of the documented 2.2864 (r7 canonical teacher)
der_student_fullset 4.8218 — same full-set harness; E4 (r7 teacher labels + Muon, vanilla ByT5-small) vs the 8.259 AdamW/r6-labels 1.0 release — a 42% error reduction at identical architecture and artifact size
parity cer_delta 0.1187pp on 600 samples
sha256 d9aa95d0e8fd0d7da80fd2f7dabedf0ac65f55ba984de565d252f5e14906d98c
license BSD-3-Clause

Runtimes reassemble split parts transparently and verify every sha256:

from secryst import Model
model = Model.load("ara-diac-small-2.0")