Skip to content

Releases: CraigVG/roman-letters-network

v2.1.0: Corpus Expansion — Cicero, Seneca, Fronto

Choose a tag to compare

@CraigVG CraigVG released this 25 Mar 01:07

Corpus Expansion

The database now spans 68 BC to 800 AD with 8,112 letters across 60 collections — up from 6,914 letters in v2.0.0.

New Collections

Collection Letters Period Source
Cicero, Ad Atticum 431 68-43 BC Latin Library + Winstedt/Gutenberg
Cicero, Ad Familiares 425 62-43 BC Latin Library + Shuckburgh/Gutenberg
Seneca, Epistulae Morales 124 62-65 AD Latin Library + Gummere (Loeb)
Fronto, Correspondence 207 139-166 AD Haines (Loeb) via Wikisource
Ignatius of Antioch 7 ~110 AD New Advent (ANF)
Cyril of Alexandria 4 430-433 AD New Advent (NPNF)

Translation Quality (continued from v2.0.0)

  • Additional 100+ fabrications found and fixed via full collection audits
  • Gregory the Great: 30 more fabrications fixed (full 399-letter audit)
  • Libanius: 55 more fabrications fixed (full 817-letter audit)
  • Augustine: 15 misaligned translations fixed (full 161-letter audit)
  • All remaining collections fully audited (1,545 letters, 14 fixes)
  • Total collections 100% audited: all 60
  • Round 5 verification: <2% fabrication rate

Stats

Metric v2.0.0 v2.1.0
Total letters 6,914 8,112
With English 6,824 8,022 (98.9%)
Collections 54 60
Date range 100-800 AD 68 BC - 800 AD
Collections audited 100% 100%

v2.0.0: Scholarly-Quality Translation Overhaul

Choose a tag to compare

@CraigVG CraigVG released this 24 Mar 22:04

What Changed

Major translation quality overhaul responding to scholarly community feedback. Every translation with Latin/Greek source text has been independently verified.

Translation Quality

  • Benchmark score: 3.54 → 4.38 (exceeds 4.0 target and 4.14 human reference)
  • 5,162 letters fully audited (100% of letters with source text)
  • ~217 fabricated translations found and fixed across all collections
  • New scholarly translation pipeline emphasizing fidelity over readability
  • 54 author-specific voice profiles for accurate register

Data Fixes

  • 3,026 AI metadata headers ("From:/To:/Date:") stripped from all translations
  • 1,628 New Advent advertising texts stripped from 19th-century translations
  • 535+ Isidore template/paraphrase translations replaced with real translations from Greek
  • 190 Hormisdas boilerplate context blocks replaced with real translations from Latin
  • 134 duplicate/misattributed records removed (99 Augustine staging, 34 Gregory duplicates, 1 Gregory IX misattribution)
  • 343 Isidore letters re-OCR'd from Patrologia Graeca page scans using Gemini Vision

UI Improvements

  • Translation source badges on every letter page (green = human, amber = AI-assisted)
  • Scholarly mode hides AI translations, shows only pre-existing scholarly sources

Infrastructure

  • scripts/scholarly_translate.py — new translation pipeline
  • scripts/benchmark_translations.py — blind comparison scoring system
  • scripts/scholarly_validate.py — automated quality checks
  • scripts/ocr_patrologia_graeca.py — Gemini Vision OCR pipeline for PG volumes
  • data/author_prompts.json — 54 author voice/style profiles
  • Deployment docs in CLAUDE.md (rsync to cvg-primary)

Database Stats

Metric v1.1.0 v2.0.0
Total letters 7,049 6,914 (duplicates removed)
With translation ~7,000 6,824 (98.7%)
Fabricated translations ~217 0
AI headers 3,026 0
Benchmark score 3.54 4.38
Collections audited 0% 100%

Audit Coverage

Every collection with Latin/Greek source text has been 100% audited:
Augustine (161), Isidore (630), Symmachus (587), Gregory (399), Cassiodorus (477), Ennodius (297), Hormisdas (249), Libanius (817), Basil (325), Pliny (365), Sidonius (171), Venantius (100), and 14 smaller collections (559).

v1.1.0: Latin/Greek Originals, Reddit Features, Performance Fixes

Choose a tag to compare

@CraigVG CraigVG released this 22 Mar 17:10

Latin/Greek coverage 40% to 75%. Manuscript scan links. Places and People pages. Scholarly mode toggle. 7 JS performance fixes. See full changelog on GitHub.

v1.0.0 — Roman Letters Network Dataset

Choose a tag to compare

@CraigVG CraigVG released this 21 Mar 03:00

Roman Letters Network Dataset v1.0.0

7,049 letters from the late Roman Empire (100-800 AD)

First public release of the complete dataset, including:

  • roman_letters.db — Full SQLite database (61 MB) with letter texts, author data, and metadata
  • letters.csv — All letters with sender, recipient, dates, locations, topics, and summaries
  • authors.csv — 1,850 historical figures with names, roles, coordinates, and biographies
  • collections.csv — 54 letter collections with source URLs and archive links
  • README.md — Dataset documentation and schema reference

Key Stats

  • 7,049 total letters from 54 collections
  • 4,410 translated to English for the first time
  • 1,850 people identified
  • Languages: English, Latin, Ancient Greek

Citation

Vander Galien, Craig. (2026). Roman Letters Network: A Digital Corpus of Late Antique Correspondence. Zenodo. https://doi.org/10.5281/zenodo.19142059

License

CC-BY 4.0