Skip to content

1.0.1: Consolidation: footguns, intent and soundness fixes

Choose a tag to compare

@lukstafi lukstafi released this 26 Aug 14:28
· 2126 commits to master since this release
894d852

Release note from CHANGES.md:

Consolidation after 1.0: making a green result mean what it says. This is
the release previously planned as v1.1, renumbered because version-number depth tracks release
scope in this project (as in the 0.6.x line): the ladder is now
1.0 → 1.0.1 → 1.0.2 → 1.1 → 1.1.1 → 1.2 (see ROADMAP.md). Its 135 closed issues were mostly
filed by PR review cycles, and their common shape is trust: a failing check that cannot be
dune promoted into a golden, an environment variable that cannot be mistyped silently, an
inlined computation that cannot lose its guard or its repetition loop, a reduction whose
accumulator width cannot depend on which schedule won, and benchmark rows that say which pass
produced them and which fixture bytes they measured. The training-loop mechanics land here too:
LR schedules, global-norm clipping, gradient accumulation, mmap-backed checkpoint loading with
its Windows arm measured and then pinned by a cross-platform regression test, and
trainable_params as distinct from "needs initialization".

Consolidation still moved the performance needle, because several soundness fixes were
performance fixes: precision-neutral accumulator localization took the Metal gpt2_mini forward
step p50 from 367.1 to 93.9 ms (−74.4%, gh-ocannl-693); batch-grid twins took the HIP step 1.72x
(gh-ocannl-643); finer fission took the tuned CUDA step 1.43x (1.56x at tf32, gh-ocannl-574);
and on CPU, whole-vector FMA took packed f32 GEBP from 12.6 to 127.2 GFLOP/s at the default
flags (gh-ocannl-614), with the auto-resolved AVX-512 width then worth 130.5 → 225.7 GFLOP/s on
a Zen 5 (gh-ocannl-621, gh-ocannl-648).

Honest nulls, recorded as such: the two-pass benchmark protocol stays OCANNL-only — no other
framework's searching cell clears ~10% on both GPU boxes (the torch.compile residue even changes
sign between them), and the old 2.5–3.5x rationale is superseded by a measured ≤10.3%
(gh-ocannl-675); the menu's descent into virtualization-inlined
scopes measured zero loops reached across the whole suite, so the provenance restriction is
insurance rather than a measured save (gh-ocannl-687); and the gh-573/gh-574 HIP ratios were
re-verified end to end after the first session's arm A routines turned out profiled-but-never-
executed in three of four cells (gh-ocannl-612).