Releases: DireLabs/ms408
Release list
ms408 v0.2.0
Everything in this release came out of an independent, AI-assisted external review of the
v0.1.0 release, plus the defects found while acting on it. Decisions are recorded in
docs/planning/i01/DECISIONS.md (L38–L40, D24–D27) and the consequences in docs/LIMITS.md.
Upgrading: evaluate()'s return shape changed and the hard-axis denominator dropped from
3 to 2. Code reading verdict["hard_axes_in_band"] or verdict["axes"] must move to
verdict["dialects"][d][...]; see Changed — BREAKING below.
Changed — BREAKING (evaluator)
- Reference bands are now stratified by Currier dialect (D21 → L38). There is no
pooled band set:reference_bands.jsoncarries one block per dialect (schema: 2) and
evaluate()returnsverdict["dialects"]["A" | "B"]plusverdict["best_match"],
instead of a single top-levelaxes/hard_axes_in_band.vms_bands(dialect)returns
one dialect's block;evaluate(tokens, dialect=…)scopes a verdict.
Why: the previous single band set was built fromA + Btruncated at the 10,000-token
budget — and Currier A alone supplies 10,709 tokens, so it contained zero Currier B
while being labelled "Currier A+B". Currier B, 68% of the manuscript, scored 0–1 of 3
hard axes against "the manuscript's" own bands. Currier A's bands and point are
byte-identical to v0.1.0; Currier B is new. zipfis demoted from a hard axis to advisory (D23 → L39). It is now unbanded,
flaggedtoken_sensitive, and excluded from the tally — so the hard-axis count is out
of 2 (h2,ed1), not 3. Why: per-dialect bands revealed that its 75%-subsample CI
is biased off the full-sample point (Currier B's own zipf point fell outside B's own
band). The fixed [10, 1000] rank window runs into the count-saturated tail at 7,500
tokens; A's bias was small enough to hide, B's was not. Same defect class asttr, and
the same existing policy is applied.
Fixed
python -m ms408.acquirefailed on a cleanpip install— the second command in the
README quickstart.RAW_ROOTassumed a repo checkout, so from a wheel it resolved outside
the package and, on system- or homebrew-managed installs, somewhere unwritable
(PermissionError). Addssources.data_home()with an explicit resolution order:
$MS408_DATA_HOME, then the repo'sdata/when running from a checkout (layout
unchanged), then$XDG_DATA_HOME/ms408(default~/.local/share/ms408).
dataset.PROCESSED_ROOTroutes through the same helper.- Degenerate token streams crashed
evaluate()with
TypeError: type NoneType doesn't define __round__— on one word repeated and on two
words alternating, the two most obvious inputs a newcomer tries.zipf_slope()documents
aNonereturn belowmin_rank + 10types, but two call sites rounded it unguarded.
Both inputs now return a normal verdict withzipf: Noneand the caveats attached. tests/test_signature.py::test_real_latin_is_excludedfailed on a clean checkout: it was
guarded on the ZL transliteration but reads an H4 corpus thatms408.acquirecannot
fetch. Now guarded on the corpus it actually needs, so it skips honestly.
Added
- The per-experiment results tier now ships (D22 → L40).
.gitignorekeeps its blanket
exclusion and allow-lists individual files thatscripts/audit_results_tier.pyclears as
metrics-only;tests/test_results_tier.pyre-runs that audit in CI, so a regenerated file
that starts embedding third-party text fails the build. Reports citing a missing
results/path fall from 19 of 28 to 12 of 28. scripts/audit_results_tier.py— L19 guard that detects both embedded running corpus text
and redistributed vocabulary slices (a JSON keyed by VMS word type is still a slice of
someone else's transliteration). Mutation-tested so it cannot pass vacuously.ms408.experiments.e34_band_dialect_scope— per-dialect band coverage diagnostic: slides
matched-budget windows across each dialect and scores them against that dialect's own
bands and the other's. Records that Currier B's bands generalise poorly within B
(ed1in band for 2 of 14 windows) — seedocs/LIMITS.md.$MS408_DATA_HOMEto override where acquired and derived data land.- Advisory axes carry their measured
subsample_biasin the artifact, so the D23 demotion
is auditable rather than asserted. ms408.verifychecks every dialect and reports cross-dialect separation asINFOrows.
Documentation
h2names its convention everywhere it is reported. The codebase computes h2 two ways
and called both "h2":textstats.char_conditional_entropy(within-word) and
textstats.lb_entropies(Lindemann–Bowern, space-inclusive, bigrams crossing word
boundaries — the convention the evaluator bands). They differ materially: 2.1247 vs 2.1643
on ZL EVA, all pages. Two live errors fixed —benchmark.pydescribed its own method as
"within-word … (Lindemann-Bowern style)", andGLOSSARY.mdquoted a within-word figure
while attributing it tolb_entropies. Slice is stated alongside convention throughout.- The Zipf rank window is stated wherever the slope is reported: least squares over ranks
[10, 1000], or to the last type if fewer,Nonebelow 20 types. It is not scale-free. docs/LIMITS.mdgains a dialect-scope section, per-dialect coverage numbers, and a note
that the "soft axes" property holds in Currier A but not uniformly in B (ungraded, pending
adversarial review — D25).
Known limitations
- 25 of 36 offline experiments cannot be regenerated from a clean checkout (D27): 22 need
the H4 corpora, whichms408.acquirecannot fetch because those sources are unregistered,
and 3 need a gitignored annotation JSONL. This caps what D22 could ship to 11 files and is
why the harness's own "real language must fail the hard axes" claim is not yet externally
checkable. - Currier B's bands are built from B's first 10,000 tokens in page order — 44% of B — and
generalise poorly to the rest of B (D24). Read a B verdict as calibrated against early B.
Full changelog: https://github.com/DireLabs/ms408/blob/v0.2.0/CHANGELOG.md
ms408 v0.1.0
MS408 is a cold, reproducible evaluator and benchmark for hypotheses about the Voynich
Manuscript (Beinecke MS 408). It does not propose a solution. It gives the research community a
shared, honest way to test an idea — "is this a cipher of Latin?", "does my generator reproduce
the manuscript?" — and to see, reproducibly, exactly where the evidence stands. Matching the
manuscript is necessary, not sufficient: an in-band result means a hypothesis is not
excluded, never that it is the mechanism.
Highlights
ms408.evaluate(tokens)+ thepython -m ms408CLI — score any word-token stream against
the manuscript's discriminator bands and get a per-axis verdict with each axis's caveat
attached (the confounded ΔI, the token-sensitive TTR, the soft mid-level syntax axes are
flagged and not counted, so the tool can't be quoted without its hedges).python -m ms408.verify— reproduce the shipped reference numbers from committed code;
--fullrebuilds the bands.- A matched-control harness (real language, cipher, and null/meaningless generators) and a
committed, firewall-built reference-band artifact. - The honest record as a feature — the archived clean-context refutation briefs document the
discipline overturning the program's own conclusions, including a cipher-exclusion headline
retracted after running a concurrently-published cipher (Greshko's Naibbe, 2025) through this
very toolkit. - No decipherment, translation, or meaning claim. Ever.
Install
pip install ms408
python -m ms408.acquire # pinned, checksummed reference data (consume-only)
python -m ms408 my_tokens.txt # ≥1000 word tokens; prints a per-axis verdictCore install needs only numpy / pandas / requests and makes no network calls on import. The
optional [vision] extra powers the annotation track.
Read more
- Docs, tutorial, and the honest-limits page: https://ms408.direlabs.com
- Methodology (harness + firewall + adversarial refutation):
docs/METHODOLOGY.md - Preprints: the constraint-envelope paper (
paper/v7) and the methods paper on adversarial
self-correction (paper/methods/v3).
Notes
- Requires Python 3.11–3.13.
- Evidence-graded (A–D) throughout; every reported number is produced by deterministic, versioned
code (the firewall). Seedocs/LIMITS.mdbefore quoting any number. - License: Apache-2.0.