Skip to content

Releases: sahrasayed3-crypto/clouda-ocr-benchmark

Clouda OCR Benchmark v0.1.0

Choose a tag to compare

@sahrasayed3-crypto sahrasayed3-crypto released this 07 Sep 16:04

Clouda OCR Benchmark v0.1.0 is the first public release of the project's reproducible distorted Arabic OCR benchmark.

Highlights:

  • 177 canonical distorted Arabic benchmark samples
  • 6 models with complete 177/177 valid coverage
  • HunyuanOCR-1.5 ranks first by normalized Arabic CER
  • Additional dots.mocr run: 177 statuses, 169 valid outputs, 8 failures
  • PaddleOCR-VL-1.6 preserved as failed/unranked evidence
  • Common 169-sample successful subset included
  • Reproducibility metadata, methodology, checksums, and public audit artifacts included
  • Dataset redistribution reviewed conservatively; assets with uncertain redistribution rights are withheld

Important:
Results are specific to this benchmark and do not imply universal OCR performance.