Skip to content

v1.0.0 (Mac)

Choose a tag to compare

@chrissab-dev chrissab-dev released this 05 Nov 12:36
· 9 commits to main since this release
333b94c

DOI 10.5281/zenodo.14336035

This release is best suited for macOS. A separate version for use in server environments has also been released.

Preprocesses images of written documents to prepare them for optical character recognition (OCR) or handwritten text recognition (HTR). Applies one of two pipelines to an image depending on its quality.

__

Aims to obtain more accurate transcriptions by making text more machine-readable. Originally written for OCR/HTR of historical documents, predominantly Tibetan-language newspaper pages from 1950-60s.

Prepares images to meet current Transkribus upload requirements then preprocesses them to be more accurately recognised by an optical character recognition (OCR) model. During preprocessing, dd_custom_preprocess applies a forked approach to images, applying one of two pipelines depending on the image quality assessment (IQA) score of a given image.

__

dd_custom_preprocess is maintained by the Divergent Discourses project (led by Robbie Barnett at SOAS and by Franz Xaver Erhard at Leipzig University. Christina Sabbagh developed dd_custom_preprocess for the project.

Full Changelog: https://github.com/Divergent-Discourses/dd_custom_preprocess/commits/v1.0.0