v1.0: Scribe Identification PoC
Features (v1.0 PoC)
Data Pipeline: A robust, multi-stage data preparation pipeline that binarizes, filters, and augments manuscript images to prevent model "cheating."
Encoder Model: A Siamese network (Encoder) trained with Triplet Loss to generate 512-dimension embedding vectors ("fingerprints") for handwriting patches.
Search Index: A high-speed FAISS index for performing efficient, large-scale similarity searches on thousands of fingerprints.
Web Application: A simple Streamlit app where you can upload a new manuscript page. The app analyzes the page, finds the most similar patches from the index, and predicts the most likely scribe based on a "voting" system.