Skip to content

7.1 release

Latest

Choose a tag to compare

@mittagessen mittagessen released this 05 Aug 16:56

What's changed

This release adds a new recognition architecture based on PP-OCRv6 with pretrained multilingual base models (tiny/small/medium). The neural architecture itself is mostly identical but hyperparameters and training setup have been optimized for historical material.

The base models are trained on a diverse corpus of historical handwritten and machine-printed materials, essentially made up of the entirety of openly accessible historical ATR datasets in addition with some private datasets and synthetic training data generated with pangoline.

The scores for all supported languages are below. It should be noted that outside of Western European languages, Hebrew, and to a lesser extent Arabic, the used datasets are usually not designed for generalized model training so their coverage is likely to be uneven and the transcription guidelines is idiosyncratic.

Language Lines Tiny CER (%) Tiny WER (%) Small CER (%) Small WER (%) Medium CER (%) Medium WER (%)
Ancient Greek 451 22.98 87.79 13.24 66.76 7.21 44.03
Arabic 1,926 20.11 68.77 13.63 52.44 9.38 38.90
Catalan 128 5.47 26.59 2.48 13.33 2.10 11.84
Church Slavonic 6,599 17.81 65.44 11.47 48.99 7.53 33.66
Classical Armenian * 3,580 0.64 3.23 0.15 0.90 0.06 0.35
Corsican 50 1.88 13.33 1.23 8.24 0.58 3.14
Czech 922 15.47 61.28 9.61 45.72 7.38 37.14
Danish 25 1.29 10.89 0.81 6.44 0.48 4.46
Dutch 4,384 14.31 51.81 8.97 35.90 6.86 27.85
English 2,005 14.43 46.77 8.18 28.45 5.90 20.35
Finnish 7,907 1.16 6.83 0.77 4.73 0.63 3.97
French 5,967 14.02 31.40 10.01 20.27 7.56 14.63
Georgian 394 28.72 85.71 19.41 70.57 12.75 51.15
German 3,007 4.61 17.96 2.51 9.86 1.62 7.20
German (shorthand) 830 41.20 84.51 25.78 62.53 17.46 47.69
Ge'ez * 2,992 1.27 5.58 0.32 1.41 0.11 0.46
Hebrew 3,426 9.11 26.32 6.50 18.48 5.06 14.21
Hungarian 38 3.20 23.05 2.82 18.79 2.40 16.67
Irish * 2,824 1.00 4.15 0.32 1.57 0.14 0.68
Italian 2,611 7.22 28.24 3.90 15.32 2.89 10.90
Latin 5,748 14.03 44.34 9.12 31.91 6.45 23.91
Latvian * 2,397 1.25 6.48 0.48 2.68 0.29 1.56
Lithuanian * 2,615 1.30 6.73 0.71 3.69 0.48 2.45
Malayalam 59 48.79 99.21 36.44 96.04 26.18 84.96
Middle Dutch 3,014 12.44 43.63 8.00 29.59 6.77 25.42
Middle French 3,970 8.70 32.38 5.03 21.14 3.41 15.46
Multilingual (mixed) 241 3.49 17.99 1.63 8.78 1.00 5.42
Norwegian 2,335 14.96 48.04 7.62 27.35 5.22 19.27
Ottoman Turkish 451 10.70 48.25 7.22 33.05 5.92 27.19
Persian 990 9.34 36.76 5.98 25.40 4.66 20.46
Polish 2,766 1.93 10.21 0.70 4.21 0.42 2.44
Portuguese 2,763 23.53 71.50 14.29 51.72 10.06 38.45
Romanian * 2,412 1.50 6.87 0.66 3.27 0.38 1.85
Russian 3,053 25.11 68.33 16.34 49.26 12.34 37.21
Serbian (Cyrillic) * 2,789 0.58 2.71 0.10 0.56 0.04 0.21
Slovak 1,550 5.11 21.50 3.24 14.08 2.62 10.70
Slovenian * 2,741 2.42 6.87 1.65 3.88 1.34 3.02
Spanish 5,465 6.19 21.49 5.00 18.82 4.49 17.61
Swedish 3,555 11.89 45.04 5.88 25.89 3.56 16.61
Syriac 1,801 10.74 44.40 6.73 30.25 4.80 21.72
Ukrainian 1,253 13.71 47.56 7.57 29.72 5.27 21.96
Urdu 1,656 10.20 41.89 5.70 25.51 3.77 16.77
Yiddish 4,320 8.96 32.79 5.81 21.54 4.49 16.24
Aggregate (micro-average) 108,010 8.71 30.00 5.43 20.22 3.91 14.98
Aggregate (macro-average) 10.99 36.16 6.93 25.33 4.93 19.07

* Evaluation uses purely synthetic data.

Another small change is how models are loaded. Instead of importing pytorch modules and trying instantiation of the weights, the code now inspects the safetensors metadata to load the require class directly by mapping it to an entry point registered during installation.