Everything in this release came out of an independent, AI-assisted external review of the
v0.1.0 release, plus the defects found while acting on it. Decisions are recorded in
docs/planning/i01/DECISIONS.md (L38–L40, D24–D27) and the consequences in docs/LIMITS.md.
Upgrading: evaluate()'s return shape changed and the hard-axis denominator dropped from
3 to 2. Code reading verdict["hard_axes_in_band"] or verdict["axes"] must move to
verdict["dialects"][d][...]; see Changed — BREAKING below.
Changed — BREAKING (evaluator)
- Reference bands are now stratified by Currier dialect (D21 → L38). There is no
pooled band set:reference_bands.jsoncarries one block per dialect (schema: 2) and
evaluate()returnsverdict["dialects"]["A" | "B"]plusverdict["best_match"],
instead of a single top-levelaxes/hard_axes_in_band.vms_bands(dialect)returns
one dialect's block;evaluate(tokens, dialect=…)scopes a verdict.
Why: the previous single band set was built fromA + Btruncated at the 10,000-token
budget — and Currier A alone supplies 10,709 tokens, so it contained zero Currier B
while being labelled "Currier A+B". Currier B, 68% of the manuscript, scored 0–1 of 3
hard axes against "the manuscript's" own bands. Currier A's bands and point are
byte-identical to v0.1.0; Currier B is new. zipfis demoted from a hard axis to advisory (D23 → L39). It is now unbanded,
flaggedtoken_sensitive, and excluded from the tally — so the hard-axis count is out
of 2 (h2,ed1), not 3. Why: per-dialect bands revealed that its 75%-subsample CI
is biased off the full-sample point (Currier B's own zipf point fell outside B's own
band). The fixed [10, 1000] rank window runs into the count-saturated tail at 7,500
tokens; A's bias was small enough to hide, B's was not. Same defect class asttr, and
the same existing policy is applied.
Fixed
python -m ms408.acquirefailed on a cleanpip install— the second command in the
README quickstart.RAW_ROOTassumed a repo checkout, so from a wheel it resolved outside
the package and, on system- or homebrew-managed installs, somewhere unwritable
(PermissionError). Addssources.data_home()with an explicit resolution order:
$MS408_DATA_HOME, then the repo'sdata/when running from a checkout (layout
unchanged), then$XDG_DATA_HOME/ms408(default~/.local/share/ms408).
dataset.PROCESSED_ROOTroutes through the same helper.- Degenerate token streams crashed
evaluate()with
TypeError: type NoneType doesn't define __round__— on one word repeated and on two
words alternating, the two most obvious inputs a newcomer tries.zipf_slope()documents
aNonereturn belowmin_rank + 10types, but two call sites rounded it unguarded.
Both inputs now return a normal verdict withzipf: Noneand the caveats attached. tests/test_signature.py::test_real_latin_is_excludedfailed on a clean checkout: it was
guarded on the ZL transliteration but reads an H4 corpus thatms408.acquirecannot
fetch. Now guarded on the corpus it actually needs, so it skips honestly.
Added
- The per-experiment results tier now ships (D22 → L40).
.gitignorekeeps its blanket
exclusion and allow-lists individual files thatscripts/audit_results_tier.pyclears as
metrics-only;tests/test_results_tier.pyre-runs that audit in CI, so a regenerated file
that starts embedding third-party text fails the build. Reports citing a missing
results/path fall from 19 of 28 to 12 of 28. scripts/audit_results_tier.py— L19 guard that detects both embedded running corpus text
and redistributed vocabulary slices (a JSON keyed by VMS word type is still a slice of
someone else's transliteration). Mutation-tested so it cannot pass vacuously.ms408.experiments.e34_band_dialect_scope— per-dialect band coverage diagnostic: slides
matched-budget windows across each dialect and scores them against that dialect's own
bands and the other's. Records that Currier B's bands generalise poorly within B
(ed1in band for 2 of 14 windows) — seedocs/LIMITS.md.$MS408_DATA_HOMEto override where acquired and derived data land.- Advisory axes carry their measured
subsample_biasin the artifact, so the D23 demotion
is auditable rather than asserted. ms408.verifychecks every dialect and reports cross-dialect separation asINFOrows.
Documentation
h2names its convention everywhere it is reported. The codebase computes h2 two ways
and called both "h2":textstats.char_conditional_entropy(within-word) and
textstats.lb_entropies(Lindemann–Bowern, space-inclusive, bigrams crossing word
boundaries — the convention the evaluator bands). They differ materially: 2.1247 vs 2.1643
on ZL EVA, all pages. Two live errors fixed —benchmark.pydescribed its own method as
"within-word … (Lindemann-Bowern style)", andGLOSSARY.mdquoted a within-word figure
while attributing it tolb_entropies. Slice is stated alongside convention throughout.- The Zipf rank window is stated wherever the slope is reported: least squares over ranks
[10, 1000], or to the last type if fewer,Nonebelow 20 types. It is not scale-free. docs/LIMITS.mdgains a dialect-scope section, per-dialect coverage numbers, and a note
that the "soft axes" property holds in Currier A but not uniformly in B (ungraded, pending
adversarial review — D25).
Known limitations
- 25 of 36 offline experiments cannot be regenerated from a clean checkout (D27): 22 need
the H4 corpora, whichms408.acquirecannot fetch because those sources are unregistered,
and 3 need a gitignored annotation JSONL. This caps what D22 could ship to 11 files and is
why the harness's own "real language must fail the hard axes" claim is not yet externally
checkable. - Currier B's bands are built from B's first 10,000 tokens in page order — 44% of B — and
generalise poorly to the rest of B (D24). Read a B verdict as calibrated against early B.
Full changelog: https://github.com/DireLabs/ms408/blob/v0.2.0/CHANGELOG.md