Preprocesses images of written documents to prepare them for optical character recognition (OCR) or handwritten text recognition (HTR). Aims to obtain more accurate transcriptions by making text more machine-readable.
Originally written for OCR/HTR of historical documents, predominantly Tibetan-language newspaper pages from 1950-60s.
Prepares images to meet current Transkribus upload requirements then preprocesses them to enhance legibility.
Includes the option to only prepare images to meet current Transkribus upload requirements.
dd_preprocess is maintained by the Divergent Discourses project (led by Robbie Barnett at SOAS and by Franz Xaver Erhard at Leipzig University. Christina Sabbagh developed dd_preprocess for the project.
Full Changelog: https://github.com/Divergent-Discourses/dd_preprocess/commits/v1.0.0