Skip to content

v1.0: Scribe Identification PoC

Choose a tag to compare

@hellcsvishal hellcsvishal released this 01 Nov 09:36
· 2 commits to main since this release

Features (v1.0 PoC)

Data Pipeline: A robust, multi-stage data preparation pipeline that binarizes, filters, and augments manuscript images to prevent model "cheating."

Encoder Model: A Siamese network (Encoder) trained with Triplet Loss to generate 512-dimension embedding vectors ("fingerprints") for handwriting patches.

Search Index: A high-speed FAISS index for performing efficient, large-scale similarity searches on thousands of fingerprints.

Web Application: A simple Streamlit app where you can upload a new manuscript page. The app analyzes the page, finds the most similar patches from the index, and predicts the most likely scribe based on a "voting" system.