A unified pipeline for processing Transkribus PageXML outputs with integrated text normalization for Tibetan text. This tool combines paragraph extraction, text normalization, and CSV merging into a single configurable workflow. The workflow integrates transkribus_XML2CSV, developed by James Engels and Christina Sabbagh, with TibNormCSV, developed by Yuki Kyogoku and Franz Xaver Erhard, and adapted by Christina Sabbagh.
Developed for the Divergent Discourses Project
SOAS University of London & Leipzig University | Funded by AHRC (UK) & DFG (Germany)
Full Changelog: v1...v2