What's changed
This release adds a new recognition architecture based on PP-OCRv6 with pretrained multilingual base models (tiny/small/medium). The neural architecture itself is mostly identical but hyperparameters and training setup have been optimized for historical material.
The base models are trained on a diverse corpus of historical handwritten and machine-printed materials, essentially made up of the entirety of openly accessible historical ATR datasets in addition with some private datasets and synthetic training data generated with pangoline.
The scores for all supported languages are below. It should be noted that outside of Western European languages, Hebrew, and to a lesser extent Arabic, the used datasets are usually not designed for generalized model training so their coverage is likely to be uneven and the transcription guidelines is idiosyncratic.
| Language | Lines | Tiny CER (%) | Tiny WER (%) | Small CER (%) | Small WER (%) | Medium CER (%) | Medium WER (%) |
|---|---|---|---|---|---|---|---|
| Ancient Greek | 451 | 22.98 | 87.79 | 13.24 | 66.76 | 7.21 | 44.03 |
| Arabic | 1,926 | 20.11 | 68.77 | 13.63 | 52.44 | 9.38 | 38.90 |
| Catalan | 128 | 5.47 | 26.59 | 2.48 | 13.33 | 2.10 | 11.84 |
| Church Slavonic | 6,599 | 17.81 | 65.44 | 11.47 | 48.99 | 7.53 | 33.66 |
| Classical Armenian * | 3,580 | 0.64 | 3.23 | 0.15 | 0.90 | 0.06 | 0.35 |
| Corsican | 50 | 1.88 | 13.33 | 1.23 | 8.24 | 0.58 | 3.14 |
| Czech | 922 | 15.47 | 61.28 | 9.61 | 45.72 | 7.38 | 37.14 |
| Danish | 25 | 1.29 | 10.89 | 0.81 | 6.44 | 0.48 | 4.46 |
| Dutch | 4,384 | 14.31 | 51.81 | 8.97 | 35.90 | 6.86 | 27.85 |
| English | 2,005 | 14.43 | 46.77 | 8.18 | 28.45 | 5.90 | 20.35 |
| Finnish | 7,907 | 1.16 | 6.83 | 0.77 | 4.73 | 0.63 | 3.97 |
| French | 5,967 | 14.02 | 31.40 | 10.01 | 20.27 | 7.56 | 14.63 |
| Georgian | 394 | 28.72 | 85.71 | 19.41 | 70.57 | 12.75 | 51.15 |
| German | 3,007 | 4.61 | 17.96 | 2.51 | 9.86 | 1.62 | 7.20 |
| German (shorthand) | 830 | 41.20 | 84.51 | 25.78 | 62.53 | 17.46 | 47.69 |
| Ge'ez * | 2,992 | 1.27 | 5.58 | 0.32 | 1.41 | 0.11 | 0.46 |
| Hebrew | 3,426 | 9.11 | 26.32 | 6.50 | 18.48 | 5.06 | 14.21 |
| Hungarian | 38 | 3.20 | 23.05 | 2.82 | 18.79 | 2.40 | 16.67 |
| Irish * | 2,824 | 1.00 | 4.15 | 0.32 | 1.57 | 0.14 | 0.68 |
| Italian | 2,611 | 7.22 | 28.24 | 3.90 | 15.32 | 2.89 | 10.90 |
| Latin | 5,748 | 14.03 | 44.34 | 9.12 | 31.91 | 6.45 | 23.91 |
| Latvian * | 2,397 | 1.25 | 6.48 | 0.48 | 2.68 | 0.29 | 1.56 |
| Lithuanian * | 2,615 | 1.30 | 6.73 | 0.71 | 3.69 | 0.48 | 2.45 |
| Malayalam | 59 | 48.79 | 99.21 | 36.44 | 96.04 | 26.18 | 84.96 |
| Middle Dutch | 3,014 | 12.44 | 43.63 | 8.00 | 29.59 | 6.77 | 25.42 |
| Middle French | 3,970 | 8.70 | 32.38 | 5.03 | 21.14 | 3.41 | 15.46 |
| Multilingual (mixed) | 241 | 3.49 | 17.99 | 1.63 | 8.78 | 1.00 | 5.42 |
| Norwegian | 2,335 | 14.96 | 48.04 | 7.62 | 27.35 | 5.22 | 19.27 |
| Ottoman Turkish | 451 | 10.70 | 48.25 | 7.22 | 33.05 | 5.92 | 27.19 |
| Persian | 990 | 9.34 | 36.76 | 5.98 | 25.40 | 4.66 | 20.46 |
| Polish | 2,766 | 1.93 | 10.21 | 0.70 | 4.21 | 0.42 | 2.44 |
| Portuguese | 2,763 | 23.53 | 71.50 | 14.29 | 51.72 | 10.06 | 38.45 |
| Romanian * | 2,412 | 1.50 | 6.87 | 0.66 | 3.27 | 0.38 | 1.85 |
| Russian | 3,053 | 25.11 | 68.33 | 16.34 | 49.26 | 12.34 | 37.21 |
| Serbian (Cyrillic) * | 2,789 | 0.58 | 2.71 | 0.10 | 0.56 | 0.04 | 0.21 |
| Slovak | 1,550 | 5.11 | 21.50 | 3.24 | 14.08 | 2.62 | 10.70 |
| Slovenian * | 2,741 | 2.42 | 6.87 | 1.65 | 3.88 | 1.34 | 3.02 |
| Spanish | 5,465 | 6.19 | 21.49 | 5.00 | 18.82 | 4.49 | 17.61 |
| Swedish | 3,555 | 11.89 | 45.04 | 5.88 | 25.89 | 3.56 | 16.61 |
| Syriac | 1,801 | 10.74 | 44.40 | 6.73 | 30.25 | 4.80 | 21.72 |
| Ukrainian | 1,253 | 13.71 | 47.56 | 7.57 | 29.72 | 5.27 | 21.96 |
| Urdu | 1,656 | 10.20 | 41.89 | 5.70 | 25.51 | 3.77 | 16.77 |
| Yiddish | 4,320 | 8.96 | 32.79 | 5.81 | 21.54 | 4.49 | 16.24 |
| Aggregate (micro-average) | 108,010 | 8.71 | 30.00 | 5.43 | 20.22 | 3.91 | 14.98 |
| Aggregate (macro-average) | 10.99 | 36.16 | 6.93 | 25.33 | 4.93 | 19.07 |
* Evaluation uses purely synthetic data.
Another small change is how models are loaded. Instead of importing pytorch modules and trying instantiation of the weights, the code now inspects the safetensors metadata to load the require class directly by mapping it to an entry point registered during installation.