0.7.1 — builds on x86_64 Linux and Windows again
0.7.0 did not build on x86_64 Linux, on Windows, or under ASan. That is
the reason this release exists; if you are on any of those, 0.7.0 is not
installable and this is the fix. Apple Silicon was unaffected, which is why
it shipped.
The rest is the test suite learning to tell a tie from a defect. Four checks
were red or blind across the three models, and every one of them turned out
to be measuring a discrete choice by the distance between logits — a
question that measurement cannot answer. None was an engine bug. The engine
changes here are two: the i8mm guard above, and one printf format.
Fixed
-
The i8mm call site is guarded on the same predicate as its source file
(docs/LEARNED.md§72).model.cdeclared and called
waste_mvq4_rows_i8mmunguarded, but the Makefile addssrc/simd_i8mm.c
toSRConly forarm%|aarch64%, so on every other platform the symbol
does not exist:model.c:805: undefined reference to `waste_mvq4_rows_i8mm'Three of five CI jobs failed at the link, and one of them is
asan + fuzz— so the sanitizers had never run against anything 0.7.0 added.
simd_i8mm.cwas already careful, defining the symbol twice so it exists
in every ARM build even where-marchdid not take; what was missing is
that a dispatcher and the list of files that satisfies it have to be
guarded on the same predicate.TK_I8MMis already unreachable off ARM,
so this is dead code being compiled, not behaviour. Verified rather than
assumed: compilingmodel.cfor x86_64 leaves the symbol undefined at
0.7.0 and unreferenced now, a full x86_64 build links, and the arm64
object still calls it. -
The two K3 checks were measuring the router, not the arithmetic
(docs/LEARNED.md§71). Chunked prefill and the CPU backend each differed
from the default path by max-abs 0.2858 against a 1e-3 threshold, and
0.7.0 shipped them red on the explanation that this was the i8mm/SMLAL
trunk kernels. It was not —trunk_kerndefaults toTK_F32, so
neither kernel was in the run.It is one routing tie. Of 1472 routing decisions in a 16-token K3 prefill,
the earliest the paths disagree on is token 12, layer 56, experts 889 and
712, at a relative margin of 7.311e-07 — the minimum over all 1472.
The 1st percentile is 4.4e-05 and the median 7.4e-03; the 47 differences
that follow average 5e-03, because they are computed on a hidden state
that has already moved. Two independent paths produce byte-identical route
traces and differ from the default in the same 48 places: NEON summation
order against scalar, disagreeing by 1e-08 on a decision that needed
1e-07.A top-K router makes an arbitrarily small arithmetic difference discrete,
so past the first flipped expert the logit distance measures how much the
model cares which of two indistinguishable experts it ran. No threshold on
it works: 1e-3 fails on a tie forever, and the 0.3 that would pass could
not catch a broken kernel.tests/route_diff.pyasks the question the threshold stood in for, over
the route and score tracesWASTE_DUMP_ROUTE/WASTE_DUMP_SCORES
already wrote: identical, tie (the first disagreement is under a
relative margin of--eps, default 1e-5) or diverged. The default
sits in the empty decade between the tie that flipped and the tightest
call that held. The three chunked/backend checks now use it, and the
result is a stricter suite, not a looser one — the argmax became a hard
failure on every path rather than a clause on one, a route that flips on a
resolvable margin fails while naming the token and layer, and a threshold
miss with the routing unchanged — the case that is a real arithmetic
defect — is its own verdict instead of being pooled with the tie. -
WASTE_DUMP_SCORESprints%.9g. The dump exists to say how close a
ranking decision was, and six digits cannot resolve one: the two scores
above both printed as0.112161while selecting differently, which reads
as a selection bug. Nine digits round-trip a float. -
The GLM oracle check was failing on both Linux jobs, on an exact tie
(docs/LEARNED.md§72). It had been red since 0.7.0, saying only "the GLM
path diverges from the oracle". linux-x86_64 and linux-arm64 report rel L2
0.00783997 and 0.00783999 — two different ISAs agreeing to five
significant figures is not a platform difference — and at the worst
logit macOS, both Linux runs and the shipped fixture all print -7.3257,
while only the freshly generated Linux oracle prints -7.43141.At layer 2, token 15 the four visible pools score 0, 0, 0 and 0.00164 with
keep=2: pool 3 wins outright and the second slot is an exact
three-way tie, margin 0.000e+00. The engine takes pool 1;
torch.topk, called inkimi_ref.pywithsorted=Falseand so free to
answer in any order, takes pool 2 on Linux and pool 1 on macOS. One pool
of four tokens attended differently is worth rel L2 0.0078 in the logits,
and it is a defect in neither.kimi_ref.pyalready wrote the trace that says this, under the same
WASTE_DUMP_DSAthe engine uses, and its own comment says why: so the two
selections "can be diffed directly rather than inferred from a logit
difference". Nothing was diffing them.tests/dsa_diff.pynow does,
reporting the same three answersroute_diff.pydoes over the pool
ranking instead of the expert ranking.Recorded because the measurement is the useful part, this is what it was
not: not a stale fixture (the engine matches it to 5.72e-06), not
macOS skipping the real comparison, not torch's version (uvresolves the
same 2.13.0, fla-core 0.5.2 and einops 0.8.2 on both), not
-ffp-contract, not thread count, not UB under ASan/UBSan, and not a
router tie — the tightest top-2-of-8 boundary gap in that prefill is
9.4e-03. -
A refusal check has three outcomes, not two.
rope_refusedgrepped
test_forward's output for the expected refusal and called anything else
"loaded instead of being refused". Atest_forwardthat is killed says
nothing, and saying nothing is not the same as loading a container it
should have rejected; the check could not tell them apart and reported the
more alarming one. It now separates them and prints what came back either
way. -
Checks say what they saw. Every comparison above now reports magnitude
and location on failure rather than a bare verdict. Two of these were
investigated for hours against a bare FAIL, and the answer each time was
four numbers the check already had in hand.