Skip to content

v1.0.0 (Server)

Latest

Choose a tag to compare

@chrissab-dev chrissab-dev released this 10 Feb 04:39
· 3 commits to server-fix since this release
a8a3fd7

DOI 10.5281/zenodo.14336035

This release is optimised for server environments. It has been adapted from the v1.0.0-mac release to ensure TensorFlow properly detects and utilises GPU acceleration.

Preprocesses images of written documents to prepare them for optical character recognition (OCR) or handwritten text recognition (HTR). Applies one of two pipelines to an image depending on its quality.

__

Aims to obtain more accurate transcriptions by making text more machine-readable. Originally written for OCR/HTR of historical documents, predominantly Tibetan-language newspaper pages from 1950-60s.

Prepares images to meet current Transkribus upload requirements then preprocesses them to be more accurately recognised by an optical character recognition (OCR) model. During preprocessing, dd_custom_preprocess applies a forked approach to images, applying one of two pipelines depending on the image quality assessment (IQA) score of a given image.

__

dd_custom_preprocess is maintained by the Divergent Discourses project (led by Robbie Barnett at SOAS and by Franz Xaver Erhard at Leipzig University. Christina Sabbagh developed dd_custom_preprocess for the project.

Full Changelog: v1.0.0...v1.0.0-server