1.0.1: Consolidation: footguns, intent and soundness fixes #788
lukstafi
announced in
Announcements
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Release note from CHANGES.md:
Consolidation after 1.0: making a green result mean what it says. This is
the release previously planned as v1.1, renumbered because version-number depth tracks release
scope in this project (as in the 0.6.x line): the ladder is now
1.0 → 1.0.1 → 1.0.2 → 1.1 → 1.1.1 → 1.2(see ROADMAP.md). Its 135 closed issues were mostlyfiled by PR review cycles, and their common shape is trust: a failing check that cannot be
dune promoted into a golden, an environment variable that cannot be mistyped silently, aninlined computation that cannot lose its guard or its repetition loop, a reduction whose
accumulator width cannot depend on which schedule won, and benchmark rows that say which pass
produced them and which fixture bytes they measured. The training-loop mechanics land here too:
LR schedules, global-norm clipping, gradient accumulation, mmap-backed checkpoint loading with
its Windows arm measured and then pinned by a cross-platform regression test, and
trainable_paramsas distinct from "needs initialization".Consolidation still moved the performance needle, because several soundness fixes were
performance fixes: precision-neutral accumulator localization took the Metal
gpt2_miniforwardstep p50 from 367.1 to 93.9 ms (−74.4%, gh-ocannl-693); batch-grid twins took the HIP step 1.72x
(gh-ocannl-643); finer fission took the tuned CUDA step 1.43x (1.56x at tf32, gh-ocannl-574);
and on CPU, whole-vector FMA took packed f32 GEBP from 12.6 to 127.2 GFLOP/s at the default
flags (gh-ocannl-614), with the auto-resolved AVX-512 width then worth 130.5 → 225.7 GFLOP/s on
a Zen 5 (gh-ocannl-621, gh-ocannl-648).
Honest nulls, recorded as such: the two-pass benchmark protocol stays OCANNL-only — no other
framework's searching cell clears ~10% on both GPU boxes (the torch.compile residue even changes
sign between them), and the old 2.5–3.5x rationale is superseded by a measured ≤10.3%
(gh-ocannl-675); the menu's descent into virtualization-inlined
scopes measured zero loops reached across the whole suite, so the provenance restriction is
insurance rather than a measured save (gh-ocannl-687); and the gh-573/gh-574 HIP ratios were
re-verified end to end after the first session's arm A routines turned out profiled-but-never-
executed in three of four cells (gh-ocannl-612).
This discussion was created from the release 1.0.1: Consolidation: footguns, intent and soundness fixes.
All reactions