Skip to content
Pummelchen edited this page Sep 20, 2026 · 139 revisions
TinyTitan Datacenter

News

What has closed, what it cost and what it taught — newest first. The Project Tracker carries only what is still open; this page is the record of everything else, because a claim without its evidence is folklore, and because most of what this project has learned came from something going wrong.

Date Item Evidence
2026-09-20 Paused here, with a rewritten handover. docs/handover-tinytitan.md described this codebase under its old identity - Pummelchen/TinyTitan, release 5.5, an eleven-install farm - and now describes TinyTitan Datacenter after 1.1: the identity and licence, the two release lineages, the current numbers for all four nodes both alone and as a chain, what the 1.1 session landed, the six open items, and the traps that have already cost time - IP-not-hostname, node4's refusal to originate, the misleading exposed_io, the derived cache budget, the two lineagues of tags, and the release sequence for a tag that may or may not be published. Its paste-in prompt is the session's starting brief. the rewritten handover
2026-09-20 All four nodes measured both ways, on the fixed binary, in one sitting. Single-node, 128 tokens at temperature 0, cache left to the default (which is 40 now that the budget is a third of physical memory): node1 7.971, node2 7.949, node3 7.902, node4 7.590 tok/s, mean 7.853 - against the release's 7.146-7.937, so the per-node figures are not merely confirmed but very slightly better, which is what the corrected default should do. Four-stage chain, same conditions: 6.143. So the cluster is still about 22% slower than a single machine, exactly the qualitative result the release published, and now reproducible from benchmark/run_four_stage_chain.sh rather than argued. The cause is unchanged and structural - the layer pipeline is serial, stage 1's 161 ms/token period being the sum of the stage times rather than the slowest of them - and the two-node reader win measured earlier (8.334 against 7.717) is still available but deliberately not taken, since docs/v4-core-design.md records the bounded reader's cost as the price of a bounded footprint. four single-node runs and three chain runs, one binary, one session
2026-09-20 The four-node chain runs again, and the released figure is confirmed rather than corrected: 6.143 tok/s. Rebuilt from PipelineWiring.swift because the wiring was recorded nowhere, validated on a two-stage pair first, then run on all four nodes with one binary (md5 f0f7c5f2), 128 tokens at temperature 0 and --expert-cache-slots 40 named on every stage: 6.206, 6.091 and 6.133 tok/s, mean 6.143, against the release's 6.016 - so the published cluster number stands within about 2%, and the per-node figures beside it were already confirmed twice. The chain rate is stage 1's period, 161 ms/token here against the 166.2 ms the release recorded: the embedder cannot start position p+1 until the sampler's token for p returns, so its decode loop measures the whole round trip, and the later stages report higher (6.371 / 6.566 / 6.753) only because their clocks start later. Stage 4 produced correct text; stage 1 prints garbage, which is correct for a ten-layer embedder. The fault that cost three attempts was mine and is now recorded: every cross-node connect must use an IP, because with a hostname connect() fails with NSPOSIXErrorDomain Code=22 \"Invalid argument\" and the stage gives up after 450 tries. benchmark/run_four_stage_chain.sh is the launch helper that never existed, so the cluster result is reproducible from the repository for the first time. three runs of the recorded script
2026-09-20 The four-node chain was rebuilt from the code, its two-stage half was validated, and the four-stage half still fails - so there is no new cluster number and 6.016 tok/s remains unverified. The wiring is not recorded anywhere (no script in benchmark/, nothing in docs/), so it was derived from PipelineWiring.swift and checked against the only surviving evidence, node logs showing node2 sending to node4 at layer 20. Validated, all four nodes on one binary (md5 f0f7c5f2): node4 listens on 47721 with STAGE_BACK_LISTEN=47712 and STAGE_BACK_ROLE=sink over layers 20:40, node2 connects to it with BACK_SWAP=1 and BACK_ROLE=both over layers 0:20, and 16 tokens at temperature 0 produced correct text - " Paris, a city renowned for its rich history, culture, and iconic landmarks.", with node2's own period 4.750 tok/s and node4's 6.671. That establishes the roles and the socket directions. The four-stage extension does not run yet: with node3 as the embedder at 0:10, node1 and node2 as middles at 10:20 and 20:30 and node4 sampling at 30:40, the head exits 1 after about 98 s with NSPOSIXErrorDomain Code=22 "Invalid argument", which points at the middle-stage back-ring hop or the forward port assignment rather than at the roles. The launch helper now exists, so the wiring is no longer unrecorded: benchmark/run_four_stage_chain.sh carries the exact commands and states in its header which half is validated. No number is claimed from it, and the release's four-stage figure stays exactly as unverifiable as it was. a working two-stage chain, and a four-stage attempt that failed
2026-09-20 The single-node target is quantified as unreachable, which closes the investigation and not the gap. The step is 141.6 ms/token split 66.5 expert I/O, 71.3 GPU and 3.8 encode. GPU plus encode alone is 75.1 ms, or 13.3 tok/s - so ~13 tok/s is the ceiling that applies only if the expert I/O becomes entirely free, with 76.9 ms of budget at the target leaving room for 1.8 ms of I/O against 66.5 today, a 97% elimination. The I/O cannot be eliminated that far on this machine: the cache is capped by RAM at 40 slots (~62 MB a slot against ~4.5 GB usable), concurrency does not help at any level (one worker matches sixteen, a serial pread beats four threads, the per-miss cost is constant across a 65% change in miss count), two independent processes split one serial SSD, and there is no batch dimension and one forward step process-wide. So the target is not reachable here by overlap or by any other measured route, by a margin of 1.8 ms - which is a finding, and the reason this goal is being closed with the number unmet rather than left open against hardware that cannot meet it. The lever that did pay was memory budget: 4.1x on the server, with the single node at 8.55 tok/s, 8.00 at release settings and 9.34 through the server. the profiler's phase split against the target's budget
2026-09-20 The one released number that may be understated cannot be checked from the record, because its command was never written down. The four-stage 6.016 tok/s figure is the release's headline cluster result, and whether it ran with the cache named explicitly cannot be answered: searching the wiki and docs/ for the stage wiring (TINYTITAN_STAGE_LISTEN, STAGE_BACK, a stage role) returns nothing. So the number is not reproducible as published - and that is a reporting defect against this project's own standard, which asks for "the commit, hardware and RAM, macOS, Swift version, exact command, exit code, complete timing footer" precisely so a result can be re-derived later. What this does and does not say: it does not show the figure is wrong, and the per-node numbers beside it are confirmed twice over (the contemporaneous tuning record says 40 slots was "the best of seven points", and a release-settings re-run gives 8.003 automatic against 7.959 explicit). It says the cluster figure is unverifiable, and that the next run of the chain must record its command and its cache size, whichever way it lands. a search of the wiki and docs for the stage wiring
2026-09-20 A claim this investigation disproved was published, and is now withdrawn in all three places it appeared. The original four-node entry above, docs/release-notes-v1.1.md, and the live GitHub release body for v1.1 all said the expert I/O was "already fully overlapped" on the strength of exposed_io reading 0.0 ms. The arithmetic says otherwise - the profiler's buckets sum to the step, and taking 42.5 ms/token of await out of a run takes 42.8 ms of step with it - so the release page was telling readers the step is GPU-bound when its I/O is on the critical path, which is the one reading that would stop someone looking where the time actually goes. The notes are corrected in place with the correction visible and a note that it was found after v1.1 shipped, and the published body was rewritten through gh release edit and verified (no line still reads "fully overlapped"). The dated entry below is left as written: it is the contemporaneous record, and the correction is newer than it in this log. Also confirmed from that entry: the single-node cache tuning found "40 slots is the best of seven points", so the per-node figures were taken with the cache tuned rather than left to a default - they stand, as the release-settings re-run independently showed. the corrected release notes, and the release body read back from the API
2026-09-20 The budget fix is confirmed neutral for the released figures, and the default now lands exactly where the release aimed. Run at the release's own settings - 128 tokens, temperature 0 - the CLI with no cache flag gives 7.998 and 8.008 tok/s, and with --expert-cache-slots 40 as released 8.002 and 7.916: 8.003 against 7.959, a 0.6% difference, so the automatic selection and the hand-passed value now agree and the fix only changes what a caller who passes nothing gets. Both sit at or above the top of the historical 7.146-7.937 range, so the per-node figures in the 1.1 release notes stand and need no correction - they were always taken with the cache named explicitly, which is precisely why the defect hid behind them. The four-stage 6.016 tok/s figure is a different matter and is still open: it is not yet established whether those stage runs named the cache size or took the default, and if they took it they were reading through a cache that paged. That is the one released number that may be understated, and it is settled by re-running the chain rather than by argument. four runs at release settings, two per configuration
2026-09-20 The page-cache inversion the cache budget cites does not reproduce, and following it would cost 21%. defaultExpertCacheBudgetBytes carries a note that its value "is only correct while expert reads bypass the cache", because under the page-cache policy "the OS holds the working set and slot memory is redundant pressure: 4-bit measured 13.61 tok/s at 16 slots against 8.78 at 128". That reads as a route to the goal, since the serial reader is the page-cache policy - so the slot count was swept against it: 8 slots 6.319 tok/s, 16 6.769, 24 7.395, 32 8.138, 40 8.552, with await falling 5564 -> 3358 ms. Monotonic, the opposite sign to the note, and 16 slots gives 6.769 rather than 13.61 - so the documented configuration loses 21% against 40 and the goal's target is not reachable that way. The note is not merely stale: it is the stated reason the budget constant exists, so the derivation above rests on a measurement that no longer holds, and the 40 slots the budget now selects is what the sweep actually supports. Recorded as a negative and as a documentation defect. Also re-verified under the corrected budget, and larger than before: the serial reader beats the parallel cache-bypassing default by 8.334 against 7.717 tok/s (+8.0%, no overlap over three runs each, await ~400 ms lower), where the same comparison read +5.1% before the budget fix. 8.552 and 8.334 are the best CLI figures this project has recorded. a five-point slot sweep under the page-cache reader, and a re-run reader A/B
2026-09-20 The budget fix reaches the CLI as well, and the earlier figures were taken with the flag passed by hand. The CLI with no cache flag now decodes 7.740 and 7.779 tok/s over 64 tokens with 3822 and 3873 ms of expert-I/O await - the same band as every explicit-40 CLI figure quoted in this investigation (7.5 to 7.9), which is the point: those numbers were always taken with --expert-cache-slots 40 given by hand, so the default path's own behaviour had never been measured and was the slow one. Nothing about the CLI changed except which cache it chooses for itself. Widths for the goal, stated plainly: ~7.76 tok/s is the CLI's own default, 9.27 to 9.34 the server's, and the target of ~13 remains unreached - but the mechanism the goal proposed for it (overlapping the read with compute) is the one that measurement ruled out, while memory budget delivered 4.1x on the server path. CLI auto against the historical explicit-40 band
2026-09-20 Fixed: the expert-cache budget is now a third of physical memory, and the server's automatic selection recovers 4.1x. affordableExpertCacheBudget clamped the tuned budget at half of physical memory, which on an 8 GB Mac mini licenses 64 slots of a 70.8 MB slot - 4.22 GiB of cache beside a ~2.3 GB floor - and the process pages. A third is the tuned value in disguise: those budgets were measured on a 24 GiB machine and a third of 24 GiB is exactly the 8 GiB defaultExpertCacheBudgetBytes, so a third reproduces the tuned number where it was tuned and scales down where a constant could not. On this model's real geometry a third is 2.67 GiB and lands on 40 slots, the measured optimum. Verified end to end: the server with no cache flag now decodes 9.27 tok/s against 2.26 before, at 4.59 GB resident with swap flat - the explicit-40 reference (9.34) within noise. Recorded trade-off: qwen38flash's 12 GiB budget is cut to 10 GiB on a 32 GiB machine; the 35B families are untouched, since their 8 GiB fits under a third of 24 GiB. The single-node target is now 1.4x away, not 2x, and the lever was memory budget rather than overlap. budget derivation, the 15-test suite, and a server run with no cache flag before and after
2026-09-20 Why the automatic selection is wrong: the budget is a constant that ignores the machine. RuntimeConfiguration.expertCacheSlots sizes the cache from defaultExpertCacheBudgetBytes, which is 8 << 30 - a fixed 8 GiB - divided by the per-slot footprint, snapped to an allowed rung. On a box with ~4.5 GB usable that budget licenses a cache which cannot fit, so the process ends up paging. This is the same failure the function's own comment documents for 8-bit on a 24 GiB machine ("64 slots against a 12 GiB budget, 34% over... cost 4.8x throughput: 0.42 tok/s at 64 against 2.01 at 32, with the hit rate falling 80% -> 67.7%, which is how a paging problem looks when it is mistaken for a cache problem") - the arithmetic was tuned against a big machine, and on an 8 GB mini it now fires for 4-bit too. Measured: the server as launched peaks at 2.14 GB resident with swap flat and decodes 2.26 tok/s, against 9.34 tok/s when the cache is pinned to 40 slots. The fix is not to lower the constant but to derive the budget from the machine - usable physical memory, not a number that assumes one - which is a change worth making with tests rather than a knob to turn. the 8 GiB constant, and server RSS and swap with automatic against explicit cache sizes
2026-09-20 The server's automatic expert-cache selection costs 4x, and it was the whole of the gap measured above. Same server, same machine, same prompt, same marginal-rate method - the only change is naming the cache size: 8 tokens in 9.20 s and 48 in 26.86 s (2.26 tok/s) with the automatic selection, against 8 in 5.48 s and 48 in 9.76 s (9.34 tok/s) with --expert-cache-slots 40. So the server was never slow: it chose a cache far smaller than the machine can hold, and the earlier entry's "about 3.5x slower than the CLI" was really "about 4x smaller cache than the CLI was given". 9.34 tok/s is the highest single-node decode rate this project has measured, above the CLI's best average of 8.112, and it re-frames the goal: the single-node target is about 1.4x away, not 2x, and the lever is cache size - the same lever the 16-to-40-slot sweep identified - not overlap. The automatic selection is the defect to pursue, since a caller who does not know to pass the flag gets a quarter of the machine's throughput, and it is worth checking whether the CLI's own default picks the same too-small value: every CLI figure quoted in this investigation passed the flag explicitly. two server runs differing only in --expert-cache-slots, measured by marginal rate
2026-09-20 Correction: the server figures above included prefill, so they are not decode rates. Dividing total request time by completion tokens counts the prompt's prefill as if it were generation, and the prompt used stopped after nine tokens, which made the error large. Measured properly - the marginal rate, from the slope between two request sizes on a prompt that does not stop early - the server decodes 40 tokens in 17.67 s, or 2.26 tok/s (8 tokens in 9.20 s against 48 in 26.86 s, same 30-token prompt). The conclusion survives in direction and changes in size: the server really is about 3.5x slower than the CLI's ~8 tok/s, but it is not 1.36 tok/s, and the concurrency scaling inferred from those contaminated numbers is withdrawn and needs re-measuring the same way. The width is not the cause - width 4 gives 1.45 and width 1 gives 1.31 by the old method, indistinguishable - so the candidate is the server's expert-cache default against the CLI's explicit 40 slots, which is where the next measurement goes. Prefill is also worth stating plainly: about 5.7 s for a 30-token prompt. marginal-rate measurement at two request sizes, and a width-4 against width-1 pair
2026-09-20 Fixed: the server could not load any model this project installs. The launch failed with "installed tokenizer is missing chat_template.jinja; reinstall the model" - and the message was wrong about the install. The tokenizer sidecar's own tokenizer_config.json carries a 7,764-character chat_template, and the code's comment claiming "the upstream tokenizer_config.json has no chat_template" is stale for these models. ServerModelSession required the sidecar file (chat_template.jinja) as well and threw when it was absent, and it then hashed that file into runtimeIdentity, so an install using the config form had no identity either. A template now resolves from either source - GFTokenizer.chatTemplateData(in:) returns the sidecar when present and otherwise the config's field, hasChatTemplate asks the question callers actually have, and the identity hashes whichever template is in force. Verified end to end: the server now starts and serves (qwen3.6-35b-a3b_4-Bit). It was never a model-install defect, so nothing was re-fetched or re-installed. First concurrency numbers from the server, and they are poor: one request 1.36 tok/s, two concurrent 0.50 tok/s aggregate - 0.37x, and far below the CLI's ~8, because the default four session slots' scratch leaves the expert cache much smaller. Together with two independent CLI processes (5.453 single against 3.326 and 3.354 concurrent - each 0.61x, aggregate only 1.23x), the conclusion is that concurrency does not pay on this hardware: adding a stream costs the existing one roughly 40%, so the SSD is a shared serial resource and the earlier hope that the idle GPU half could absorb a second sequence is refuted. server launch before and after the fix; two-process and two-request concurrency measurements
2026-09-20 The concurrency measurement is blocked by the model install, not by the engine. With the server binary deployed and the width flags confirmed real (--max-concurrent-sequences, 1 to 4), the server refuses to load the model that is installed on all four nodes: error: installed tokenizer is missing chat_template.jinja; reinstall the model. The CLI runs that same install happily because a raw prompt needs no chat template, which is why this has never surfaced here. It is reported rather than repaired: the project's own rule is that a check which cannot run is reported not checked and never fixed by downloading, converting or re-installing a model, so the four-sequence aggregate prediction from the previous entry stays unmeasured and is not to be counted as a result. What is needed is a model install that carries its chat template - an operator job - after which the two predictions (about 28 tok/s aggregate if steps time-slice, about 7 if they serialise harder) can be settled in one run. server launch on a node, refusing to load the model
2026-09-20 The server is already built at width four; only the step gate stops the speed-up. ServerArguments defaults maxConcurrentSequences to 4 (--max-concurrent-sequences, 1 to 4) and derives sessionSlots from it, documented as "the width a session's runner and scratch may actually be built with" - so the runner and its scratch are not single-slot, and four sequences can be admitted and held at once. What ForwardStepGate prevents is only that their forward passes interleave. That gives the next experiment a sharp prediction and makes it cheap: because the step leaves the GPU idle for 66.5 of its 141.6 ms, four admitted sequences running one step at a time time-sliced should already reach about 4 x 7.06 = 28 tok/s aggregate on a single node with no new code - throughput that until now has been credited to four separate machines. True batching is what would raise the rate per sequence (two sequences pipelined per layer give max(3.3 ms of reads, 3.56 ms of compute) per layer, so ~14 tok/s each), but the aggregate half of the prize appears to be already present and simply unmeasured. ServerArguments.swift defaults and the ForwardStepGate finding
2026-09-20 The overlap the goal needs is gated off in the engine, on purpose, and there is no batch dimension to put it in. The runtime has no batchSize, sequenceCount or decodeBatch anywhere; RealForwardRunner holds no array of sequences. The server can admit several generations at once - ServerCoordinator(queueLimit:width:), width defaulting to 1 - but its own comment says what happens next: "The engine's ForwardStepGate keeps their forward passes from interleaving." And ForwardStepGate is exactly that: an actor with a busy flag and a FIFO waiter list whose acquire() waits "until no step is in flight". So at most one forward step exists process-wide, and one generation's expert read can never run while another's compute does - which is the only remaining overlap that reaches the target, and it explains every negative this investigation produced: the read and the compute are serial because the engine is a serial single-stream decoder. The route to ~14 tok/s is therefore two sequences in flight, pipelined per layer - issue the next token's read alongside this token's compute (2 reads 3.3 ms under 2 computes 3.56 ms per layer, so max not sum, 1.78 ms/token). It needs a batch dimension or a per-sequence step gate, not a faster disk. the runtime's inference sources and ForwardStepGate.swift
2026-09-20 The serial reader's win is not the page cache, and the device does not reward concurrency. If reading through the page cache were what makes the serial fallback faster, repeated runs would improve as the working set warmed. They do not: three serial runs in a row give 7.998, 8.042, 8.036 tok/s and three parallel/bypass runs 7.832, 7.793, 7.745 - both flat, the serial path simply about 3-5% ahead every time. So the difference is the reader implementation, not cache warmth. And a serial reader beating a four-thread one is the third independent sign that this device does not reward concurrency: one I/O worker matches sixteen, four reader threads lose to one plain pread loop, and the per-miss cost is unchanged by a 65% change in miss count. Three measurements, one reading - the expert read is a fixed serial latency, so the only route left to a shorter step is to need fewer reads, not to run them wider. six consecutive runs, three per reader
2026-09-20 The default expert reader is the slow one: the parallel, cache-bypassing reader costs 5% against the plain serial fallback. PreadExpertStreamer picks a ParallelExpertReader(threads: 4, bypassCache: true) by default and falls back to a serial readFull loop only when TINYTITAN_BOUNDED_IO=0. Alternating the two, three runs each: serial 8.101, 8.102, 8.133 tok/s against parallel 7.687, 7.744, 7.729 - 8.112 against 7.720, or +5.1%, with no overlap between the groups, and expert-I/O await down from 1921-1962 to 1824-1836 ms. Bypassing the page cache makes every read go to the device; reading through it lets the working set stay warm. Also measured and negative: the I/O worker count is not the constraint. ExpertIOScheduler was hardcoded to 4 workers and is now TINYTITAN_EXPERT_IO_WORKERS; one worker performs exactly as well as sixteen (7.498 / 7.470 / 7.486 / 7.520 tok/s at 1 / 4 / 8 / 16), which proves the reads were never running concurrently in the first place - so the serialisation is below the scheduler, and adding concurrency there cannot help. alternating reader A/B, three runs each; worker-count sweep of four
2026-09-20 The engine already splits hits from misses, so the remaining question is narrower than 'overlap the read'. Reading the decode path: earlyHits keeps a phase1HitCB command buffer that "already exists and is committed", hits and misses are carried in separate slot lists, and the miss fetch is submitted early (beginFetchRoutedExperts, event-driven when enabled). So the structure to compute resident experts while the one missing expert loads is already there - and it is not paying, which is the finding. The arithmetic says why it cannot pay as built: await divided by the per-read cost is a constant miss count - 108.6 ms/token at 16 cache slots against 66.1 at 40 is 66 versus 40 misses, at ~1.65 ms each. Reads cost their full latency individually and do not cover each other, and a layer has only ~1 miss, so there is nothing to batch. The untested direction is more lead time rather than a split: prefetchAhead accepts only 1 or 2 and the ring holds about 2 slots (free=0.98), so the engine can never have more than two reads outstanding, however far ahead it predicts - and a deeper ring was measured worse with a deeper prediction. Depth and slots have never been raised together. the decode path in RealForwardRunner+Decode, and the cache-slot measurements
2026-09-20 The expert cache is capped by RAM, so a larger one cannot reach the target either. Peak resident set against cache slots, one node, 16 tokens: 16 slots 3.33 GB, 40 slots 4.83 GB, 48 slots 4.97 GB, 64 slots 5.14 GB. That is ~62 MB per slot, and 40 slots already peaks at 4.83 GB on a machine whose usable memory is about 4.5 GB. Swap confirms the wall rather than merely suggesting it: flat at 40 (775 M), rising at 48 (775 -> 855 M) and exploding at 64 (855 -> 1610 M) - and throughput follows it down, 7.289 -> 6.989 -> 5.576 tok/s. So the non-cache floor is about 2.3 GB, which leaves about 2.2 GB for experts, and that is the 40-slot optimum by arithmetic. Even freeing a generous 1.15 GB would buy only ~18 more slots, worth perhaps 8 tok/s and not 13. The cache route to the goal is therefore closed, and for the same reason as everything else: on one 8 GB Mac mini, the read and the compute are one serial path and there is no memory left to shorten the read with. What is left is a second token in flight, which is the only structure that overlaps them. four profiled runs with peak RSS and swap deltas
2026-09-20 The knob surface is exhausted, and one knob proved the two big phases are a single critical path. Five profiled runs on one node: baseline 7.397 tok/s, EARLY_HITS=1 7.418, PREFETCH_IO_TIER=standard 7.450, the two together 7.430, HC_FUSED=1 7.414 - a 0.75% spread, which is run-to-run noise, so none of them is a lever. But EARLY_HITS moved time between the buckets: expert-I/O await rose 2069 -> 2172 ms while unaccounted GPU waits fell 2128 -> 2001 ms, and throughput did not change. A knob that trades one bucket for the other, one for one, with no net effect is evidence that they are not two independent costs but one serial critical path - which is what the cache sweep already said (delta-await 42.5 ms against delta-step 42.8 ms). So the per-layer structure is read-then-compute, 1.65 ms and 1.78 ms per layer, and the sum is the step. Overlapping them needs a second token in flight; shrinking one is the only other route, and the only one with measured headroom is the expert cache - the await falls 108.6 -> 66.1 ms/token from 16 to 40 slots, and the machine swaps past 64. five profiled runs on one node
2026-09-20 The prefetch ring is worth 15%, coverage beats precision, and precision actively hurts. The ring's two counters were only ever read by the server, so the CLI could not say whether it did anything; both are printed now. On against off, all else equal: 7.385 and 7.437 tok/s with the ring, 6.357 and 6.428 without - worth about a tok/s, and it removes ~600 ms of the 2660 ms of expert-I/O await. But the ring adopts only 55% of what it issues (1003 issued, 553 adopted), and the engine's own comment says a wrong read costs the next layer's demand reads their device time - so the precision margin was swept: it raises precision to 78% and loses 0.56 tok/s, with the await rising from 2097 to 2376 ms. So the comment is wrong and the intuition is backwards: a wasted prefetch is nearly free, a missing one costs a full blocking read, and the ring is coverage-limited rather than precision-limited. That changes what to fix: issue more, earlier - not better. exposed_io still reports 0.0 ms while 64 ms/token of await sits on the critical path. four profiled runs on one node at margin 0, 0.02, 0.03 and 0.06, plus ring-on/ring-off pairs
2026-09-20 The expert read is on the critical path, and the instrument that said otherwise is wrong. Two runs on one node, same prompt and 32 tokens, profiler on, cache size the only change: at 16 slots the step is 184.4 ms/token with 108.6 ms/token of expert-I/O await, and at 40 slots it is 141.6 ms/token with 66.1 ms. The deltas are -42.8 ms and -42.5 ms — a ratio of 1.007, so removing expert-I/O time removed step time one for one and the read is fully serialised. That contradicts totalExposedIoNanos, which reports exposed_io = 0.0 ms in both runs and therefore cannot be answering 'does this I/O add to the step'. The earlier conclusion recorded here - that the overlap already worked and the critical path was GPU - is withdrawn, and the arithmetic had already shown it: the profiler's buckets sum exactly to the step (66.5 + 71.3 + 3.8 = 141.6), and buckets that sum are buckets that do not overlap. The prize is therefore real and it is on one machine: 66.1 of the 141.6 ms is I/O, so if it could be hidden behind the 71.3 ms of GPU work the step approaches 75 ms and ~13.3 tok/s, about 1.9x today's 7.07 - larger than anything the four-node work has produced, and it needs no network. two profiled runs on one node at 16 and 40 cache slots
2026-09-20 Wiki review: the dependency framing is gone, and the plan pages are flagged as stale. Every non-historical reference to a 'sister project' has been removed - Home's 'what it builds on', the Glossary entry, two Architecture rows, the Roadmap objective, the tracked-release and adopt-their-kernels rows in the tracker, and Testbed's 'the sister project as a dependency' section - because ttd is built from scratch and is MIT, so there is no upstream to track or adopt from. What is left is a plan that describes a different project: Roadmap.md, the milestone pages and most of the tracker's rows still describe the retired DatacenterEngine line (M0-M5, expert-parallel sharding and the all-reduce that goes with it), while main ships the streaming runtime and the layer-pipeline chain measured at 7.1-7.9 tok/s per node and 6.0 through four. Architecture.md's sharding section says expert parallelism; what shipped divides layers. Rewriting those for ttd is a planning decision rather than a cleanup, so it is flagged here rather than done quietly. Dated News entries naming the sister project are history and are left as they stand. commits on the wiki
2026-09-20 ttd is not based on turbo-fieldfare, and the tree no longer says it is. The project is built from scratch and is MIT. RELEASE.md had stated outright that the repository was a GitHub fork of that project and deliberately left as one; that section is replaced by the standalone standing, and the same claim is corrected in README.md's Credits, CONTRIBUTING.md's --repo rule, tools/release.sh's header comment, the security-advisory link and the site's lineage paragraph. Earlier entries on this page describe taking code from the sister project; those are dated records of what was done then and are left as they stand. commits on main
2026-09-20 TinyTitan Datacenter is ttd, and it is a dedicated repository. Short name for the project: ttd. It stands on its own - it builds and runs with no other repository, and its installer, release target, plugin metadata, site documentation and wiki no longer point at the runtime it was originally taken from. The Apache-2.0 attribution for that source is kept in NOTICE and THIRD_PARTY_NOTICES.md, which is a licence obligation rather than a link. the front page and instruction file; commits on main
2026-09-20 main now carries the engine that was actually measured, and the branch list is down to one. The repository's main is the streaming runtime that produced the four single-node benchmarks - 7.146 to 7.937 tok/s, mean 7.73, one binary, 128 tokens, every machine above the reference's best of 7.075 - and the four-stage cluster chain at 6.016 tok/s. It is 28 SwiftPM targets under sources/TinyTitan*, built on the reference's single-node runtime with this repository's distribution on top, which is the direction D123 set. The previous main is retired rather than deleted: the DatacenterEngine line at 55d84fa, with its decision records through D457, is preserved as the annotated tag retired-datacenter-engine, restorable with git checkout -b datacenter-engine retired-datacenter-engine. The distribution branch has been deleted - it held the same engine's earlier state, whose tip was an ancestor of the new main 46 commits back - and only main remains, locally and on the remote. One consequence worth stating: the two lines are separate projects with separate licences, MIT on the retired side and Apache-2.0 on the engine, so the replacement is a change of what this repository ships rather than a merge of the two. the retired tag retired-datacenter-engine; the four-stage chain and the four single-node benchmarks recorded above; D436 to D457 in that tag's docs/repository-decisions.md
2026-09-20 The port reaches 7.1-7.9 tok/s on every one of the four Mac minis, and the four-node cluster is still slower than a single machine - which is now a measured result rather than a suspicion. Four single-node runs, one at a time, one binary (md5 ec6710cd... on all four), 128 tokens at temperature 0: node1 7.935, node2 7.937, node3 7.884, node4 7.146 / 7.448 / 7.396 tok/s, mean 7.73, all four producing the same text. The project's own engine began at 0.108 tok/s cached at M1 and had reached ~1.6 tok/s by D117; the ported streaming runtime now measures 7.1-7.9, and every machine beats the reference's best published single-node figure of 7.075 - its 5.164 / 6.019 / 7.075 at 1 / 2 / 3 GB of expert cache, and 2.756 at 4 GB where it swaps. The cluster does not follow. A four-stage layer pipeline runs the model correctly on all four machines and measures 6.016 tok/s at 128 tokens, below one node, and the reason is measured rather than argued: the head's period is 166.2 ms/token, which sits at the serial prediction (the sum of the stage times, 194.8 ms) and nowhere near the pipelined one (the slowest stage, 48.7 ms), because stage 1's step for position p+1 needs the token the head samples during its step for p, so the chain can only alternate. A layer pipeline divides memory, not time. The step itself is 47% expert I/O and 50% GPU wait, and the I/O is already fully overlapped - totalExposedIoNanos, an instrument the engine maintained and never printed, reads 0.0 ms on all four nodes. Every single-node tuning axis is at its optimum: prefetch depth, four tunables, prefetch-ring depth, and the expert cache itself, where 40 slots is the best of seven points and 64 and above collapse into swap (7.4 to 4.9 to 2.6 tok/s). Reaching 21 tok/s needs the wire rather than the engine: the farm's only active link is 1 GbE (118 MB/s, 765 us round trip, and UDP measures the same as TCP so it is the link and not the protocol), the Thunderbolt and 10GbE ports report status: inactive on all four machines, pooling expert residency across peers is 8.6x slower than local disk, and tensor parallelism would spend 61 ms of synchronisation per token against a 47.6 ms target. D436-D457 in docs/repository-decisions.md; the four-stage chain and the four single-node benchmarks; docs/distribution-design.md
2026-09-18 The decoded-layer cache was re-tested at the new operating point, and it still loses. D89 rejected it when a step was 5.42 s; nine rounds later the step is 0.626 s and load is 68 ms, so the trade was worth re-asking. 1 GB takes load 68 -> 20 ms — 70% of the phase — and the step is still 2% worse (0.626 s against 0.639, three alternated pairs), because mix.gather absorbs 16 ms of it. The default stays 0: the fourth independent confirmation of the memory wall D118, DC-117
2026-09-18 The bf16 tile was loaded one instruction per row — D110 in a different kernel. The head fills a 32x32 tile with one warp-wide 64-byte load per row: 32 instructions for 2 KB, and head measured 12 GB/s on hardware whose memory does ~100. Four bf16 per lane with ushort4 — eight loads instead of thirty-two — took head 82 -> 48 ms and the step to 0.628 s, 1.591 tok/s, bit-identical by construction because the tile's contents are unchanged. A load can be perfectly coalesced at the warp level and still instruction-bound; the way to see it is bytes-per-second against the hardware D117, DC-117
2026-09-18 The dense projections were copied to the device 130 times a step. D115 fixed that for the head; the same defect sat in the dense path, 865 MB of copying a step for weights that never change. packedTensor hands over the whole tensor, so one bytesNoCopy mapping covers all three sections and each dispatch addresses them by offset: attn.core 146 -> 118 ms, step 0.678 -> 0.65-0.67 s (about 1.5 tok/s). An expert's row range must never be mapped — those bytes come from the evictable slab cache — and the key is content-addressed (name#sha256-prefix), because a name is unique within one install and says nothing across two D116, DC-117
2026-09-18 The head was copied to the device every step, and the read wall is now measured from four sides. MetalBf16Matmul.matmul copies its whole weight per call, and the head is 1.017 GB in 31 blocks — 1.04 GB of copying a step for ~1 ms of arithmetic. Mapped once with makeBuffer(bytesNoCopy:) over the payload the engine already holds: head 116 -> 82 ms, +18% on the step (0.757 s against 0.896 in four alternated pairs), default 0.678 s, 1.476 tok/s. The key must name the (tensor, row window) pair — on the name alone node 1 reused node 0's head and the sharded test produced the wrong tokens. And the read: the device's cold sequential rate is 1.65 GB/s and the preload already beats it (1.82 cold, 20 GB/s warm); the 582 MB/step of slabs is reused cross-step (43% at 1 GiB, 53% at 1.5) but the memory costs more than it saves in every configuration tried, including with the page cache disabled. 7 tok/s is 143 ms/step against a 238 ms read and 440 ms of everything else D115, DC-117
2026-09-18 A layer's experts are independent, so they are asked for in one dispatch. Each of the 640 fused calls a decode step made built a command buffer, committed it and blocked on it, and at these shapes the kernel is microseconds of a ~0.3 ms call — the wait was the cost. One command buffer per layer instead of one wait per expert: mix.read 179 -> 84 ms, mix.down 151 -> 52 ms, step 0.919 -> 0.722 s, 1.384 tok/s. The batch is taken only when every chosen expert was routed the same number of rows (every decode step, not every prefill). The bug worth remembering: the batch's 40-buffer request shared a grow-only buffer cache with the dense projections' 5-buffer request, and that cache reallocates its whole set when the count changes — 2.47 s/step against 0.92, showing in mix.gather, a phase the change never touched D114, DC-117
2026-09-18 The preload was fanning out over half the work, and a partial payload cache is catastrophic. A slab is three preads, so fanning over eight experts gave eight threads six sequential reads each; one task per (expert, projection) is the same bytes with twice the requests in flight, and five runs moved the median 0.984 -> 0.919 s (1.016 -> 1.089 tok/s). The warning is the other measurement: 256 MiB of payload cache gives 2.34 s/step and 1024 gives 2.58, because the LRU holds part of the dense payload and not the head, so both re-read every step — zero is better than either (1.359 s) because it streams uniformly. 2048 MiB is near-minimal and load-bearing, not a knob D113, DC-117
2026-09-18 The attention projections were the same change, and the slab cache got smaller once the kernel cached too. The four int4 attention projections are 1.02 GB of fp32 a step and now multiply from their stored form like the GDN three: load 119 -> 54 ms, step 1.087 -> 1.025 s. And with the buffer cache on (D111), the slab cache's optimum moved down — three alternated pairs put 128 MiB at 0.930 s against 256 at 0.957 — while zero is worse than every non-zero size (1.595 s), because preloadPacked declines with no cache and the read fan-out disappears. Default 128 MiB. 0.920 -> 1.016 tok/s, peak RSS 2.57-2.67 GB. 63% of the step is now the expert path D112, DC-117, DC-127
2026-09-18 The page cache is worth having, and the dense int4 projections should never have been decoded. The install can be read through the kernel's buffer cache (SHARD_INSTALL_CACHED, on by default) — right for a repeating 582 MB working set even though it was wrong for the 20 GB sequential scan D58 records: cold mix.gather 469 ms, warm 266-286, with RSS, free disk and swap unchanged. And the three Gated DeltaNet int4 projections (581 MB of the install's 738 MB of dense int4) now come from their stored form through the fused kernel, so the fp32 array is never built: load 350 -> 119 ms, digest unchanged. The bug between them is the one to remember — the first packedTensor used the row-range reader and sent 7.5 GB a step back to the device for bytes the payload cache already held, visible only in the byte counter (reads 8.46 -> 14.28 GB). 0.682 -> 0.920 tok/s D111, DC-117, DC-127
2026-09-18 The int4 kernel was loading one byte at a time — fixing that made the packed expert path pay. The fused expert kernel measured 4.4 GB/s on memory that does ~100: lane l and lane l+1 walked rows 1 KB apart and every element was its own uchar load, so each fetch used one byte of a cache line. Reading each row as a uint4 (32 codes) made it 3.6-4x faster (32768x2048: 8.734 -> 2.196 ms), and the grid caught a trap on the way: #pragma unroll let .relaxed reassociate the accumulator chain and moved one output by 1 ULP, so the unroll is gone and the pipeline is .safe. The packed slab cache was swept and smaller is better — 256 MiB 1.465 s, 512 1.467, 768 1.483, 1024 1.667 — because the resident bytes cost more in memory pressure than the reads they save (D106 a third time). Five alternated pairs: fused 1.465 s against the split path's 1.539, digest unchanged, at lower peak RSS. Both defaults are now the measured ones, and three suite-found bugs are fixed rather than excused: a 256-byte default budget that made the cache look empty, a preload that chose its destination by switch instead of capability, and a fused counter that reported 13,824 elements read on a forward served entirely from memory. 0.638 -> 0.682 tok/s D110, DC-117, DC-127
2026-09-18 The LM head comes off the device every step — and three of the round's ideas were losses, which is what bounds the rest. A fresh profile put the decode step at 1769 ms (mix.read 805, load 361, head 255, attn.core 202) and 0.565 tok/s. Lost and recorded as lost: the GPU unpack for the dense tensor(named:) path (load 353 → 534 ms); the fused int4 expert path, which is a wash (2.29 s/step with no slab cache, 1.575 with a 1 GB one against a 1.69 s control — 640 per-slab dispatch-and-wait pairs eat what the kernel saves); and the first head-residency attempt, which took the node into D106's swap failure twice (31 concurrent blocks each reading the whole 1.017 GB, and UncachedFile.readData building every read twice — 2.03 GB transient for one 1.017 GB block). Both bugs are fixed. Kept: the LM head on the GPU from its stored bf16 (MetalBf16Matmul) — head 255 → 118 ms, step 1.718 → 1.580 s in an alternated A/B, digest unchanged — with three small measured fixes (per-slot grow-only MetalBufferCache, the GDN per-channel allocation hoisted out of an 8,192-iteration loop, the head's blocks on the serial matmul body). Default now 1.567 s/step, 0.638 tok/s, peak RSS 3.10 GB, 266 tests, 0 failures. The study of the sister project says what is left is not arithmetic: ~1.5-1.8 GB of dense plus ~0.17 GB of experts per token, all packed and never widened to fp32, its head int4 at 286 MB against this install's bf16 at 1017 MB, and ~55 ms of its 141 ms is host blocked on the routing readback and the miss reads D109, DC-117, DC-113, DC-127
2026-09-18 The fused int4 matmul works on the GPU — 1.2–2.4× faster — and it is still not the lever: the unpack was never the bottleneck. MetalInt4Matmul dequantises inside the matmul, so the ~6.5 GB of Float a token materialises never exists at all. It is bit-identical to InstallFile.dequantizeInt4 followed by Ops.orderedMatmul over the D107 grid, and in release it beats unpack-then-matmul 1.23× / 1.45× / 2.36× on the real gate_up, down and head-width shapes — 0.114 s of a 1.74 s step, about 6.6%, the same order as the ~7% the CPU probe bounded. mix.read is the disk read, so the route to 7 tok/s is residency, and this kernel is its precondition rather than its win. The grid also found something that had been assumed away: an Apple GPU flushes a denormal product to zero where the CPU keeps it, for this kernel and for the MetalMatmul already in the tree, and no math mode changes it — now a named test instead of an implicit assumption. Kept as a tested, measured asset, not wired in yet D108, DC-113, DC-127, MetalInt4MatmulTests (6)
2026-09-18 The fused int4 kernel is 4.7× slower than the split it was meant to replace — and the CPU phase is over. DC-120's idea was to dequantise in the matmul's inner loop so the fp32 slab never exists (mix.read writes 6.5 GB of Float per token and reads it straight back). The kernel was built and proved bit-identical to dequantise-then-multiply over a grid of shapes and the model's real ones, so only speed was open — and in release the split path takes 0.33 ms (unpack 0.23 + matmul 0.10) against 1.55 ms fused: 4.7× slower (0.14× in debug). Dequantising per element destroys the vectorisation the eight-wide unpack and the four-wide matmul each enjoy. The probe bounded the prize too: a 512×2048 slab unpacks in 0.23 ms, so a token's ~520 slabs cost ~0.12 s, ~7% — there was almost nothing to win. Reverted, not kept, on the DC-122 precedent: the number is worth more than the code. With D105 (the read is saturated), D106 (the cache is unaffordable) and now this, every CPU-side lever is measured and closed — the engine's split is the right shape for a CPU, and the distance to 7 tok/s is the format and the device D107, DC-113, DC-120
2026-09-18 A packed slab cache: the mechanism works, the node cannot afford it — and the pattern is now unmistakable. D105 said the device is saturated, so the lever is reading fewer bytes; InstallFile.SlabCache holds expert row ranges as stored (four bits per weight, not the thirty-two the decoder makes) and a hit skips three preads. Measured: 512 MB never hits (a slab is evicted before the next token asks); 768 MB hits 36.1% and reads 0.73 GB/step; 1024 MB hits 45.0%, reads 0.65 GB/step instead of 1.08, and mix.read falls 0.80 → 0.54 s — and the step gets worse (1.931-2.022 against 1.735), because load nearly doubles (0.363 → 0.591-0.650) and head grows (0.261 → 0.411) on a machine swapping 1.94 of 3.07 GB. It is the fourth time this engine has reached that verdict (D89, D98, D105, now D106), and the pattern is the finding: this node's binding constraint is memory — buying a saturated device with RAM costs more elsewhere than it saves. Default 0, SHARD_SLAB_CACHE_MB is the knob, digest identical in every arm D106, DC-120, DC-126, Install.swift, SlabCacheTests.swift
2026-09-18 DC-121 is built, measured, and it loses — and the loss is the useful part. The reference's prefetch ring is a good idea and the reasoning held: a layer's routing is correlated between consecutive tokens, and D101 had left the expert reads on the critical path, so issuing the last token's choices early on a background thread before the attention should hide them. It did not. Alternated on one binary: 1.859 / 1.816 s/step with it on against 1.854 / 1.721 with it off, with mix.read larger when it is on (0.830 / 0.813 against 0.790 / 0.760) and the digest identical throughout. The reading: D101's fan-out had already taken the latency out of the read, so the device is now saturated and a second stream of reads does not fill a gap — it takes bandwidth from the first. SHARD_EXPERT_PREDICT=1 turns it on; the default is off, because a slower thing is not a default, and the mechanism has a force hook so it stays tested. The lever therefore moves from hiding reads to reading fewer bytes: the bank holds [Float] today, so 512 MB buys ~61 slices of the ~773 a token asks for, and the same bytes packed as int4 buy four times as many D105, DC-120, DC-121, ExpertProvider.swift
2026-09-18 The contract matmul is threaded, and this time it keeps the contract: 0.443 → ~0.55 tok/s. The first attempt (DC-122) moved bits and was reverted; the rule it was missing is now written down and tested. The serial body's four-wide groups are the fixed partition {0-3}, {4-7}, … and its scalar tail is exactly the last out % 4 columns, so a chunk ending at a non-multiple of four re-cuts the grouping and moves the tail. With every chunk boundary a multiple of four, a non-final chunk covers whole groups and reaches its end with no tail at all, and the final chunk ends at out and performs the serial body's own tail — same additions, same order, only the thread differs, which is D63. Measured on the phases, alternated on one binary: attn.core 0.493/0.495 → 0.209/0.202 (2.4x), mix.gateup 0.224/0.225 → 0.077/0.073 (3.0x), mix.down 0.102 → 0.049/0.047, step 4.124/4.103 → 1.866/1.797 s, digest identical in all four arms. OrderedMatmulThreadTests walks the shapes a regrouping would show in: out % 4 ≠ 0, out < 4, k = 1, rows > 1, thread counts 0/1/2/3/5/8/16 and the model's real projection and expert shapes D104, DC-122, Ops.swift, OrderedMatmulThreadTests.swift
2026-09-18 DC-124 is found and fixed: a data race in the checkpoint reader's streaming handle. The intermittent abort that had twice turned the Python gate red was Safetensors.StreamHandle.file, a bare var read and written from the reader's worker threads: two workers see nil together, each opens a descriptor, and the racing assignments release one UncachedFile twice — the Swift runtime's "deallocated with non-zero retain count … a dangling reference", then a SIGABRT. The discriminator that pointed at it was that every failure was on a checkpoint path and every install run was clean: InstallFile holds its handle in a let. The handle now opens once under a lock. 3 of 4 pair runs failed before, 0 of 5 after, and the retain-count line appears zero times — five clean runs by luck against a 75% prior is about 0.1%, which is the evidence this rests on; DC-125 tracks turning it into a deterministic test. Also in this round: the expert adapter's metric writes are guarded, since D101's fan-out writes them from several threads DC-124, DC-125, Safetensors.swift, ExpertProvider.swift
2026-09-18 The head's vocabulary blocks fan out — 0.400 → 0.443 tok/s, and the safe half of DC-122 is done. DC-122 threaded the general contract matmul once and it moved bits (bounding each thread's work by its own last regrouped the four-wide columns and moved the scalar tail), so that function is the original single-threaded body with the attempt on the record. The head is a different kind of problem: the unit of work is a whole vocabulary block of rows, so every block runs the same Ops.orderedMatmul call with the same internal grouping and only the thread differs — which is exactly what D63 permits. Measured, alternated, one binary: the head phase 0.434 / 0.439 → 0.259 / 0.255 s/step, digest identical in all four arms. HeadLogitsTests walks block widths 1, 3, 7, 64 and 512 and compares bit patterns — deliberately including out % 4 != 0, the regime that caught the earlier attempt — plus a slice test and one that proves a failed block is raised rather than swallowed into zeros. 0.443 tok/s, and the general matmul's turn is still open with its menu of decompositions in DC-122 D102, DC-122, ModelCache.swift, HeadLogitsTests.swift
2026-09-18 DC-124: an intermittent abort, seen twice, never reproduced. Two consecutive full gate runs produced two different Python failures that were the same event — a child process terminating abnormally against the gate's declared 4.20 GB with swap at 1.21 of 2.05 GB used. Isolating it: the m1 gate test then passed 3/3 (fan-out on, off, on), the sharded suite 5/5 alone, and datacenter-trace 3/3 — so it is not the D101 fan-out, and the third full run is green. The Swift runtime's "UncachedFile deallocated with non-zero retain count" line is very likely teardown noise for an abort that happened elsewhere; that class is a plain deinit { close(descriptor) }. Recorded as a row rather than a footnote because it is unexplained, with a done-when that requires reproducing it and capturing the abort's own message DC-124, D101
2026-09-18 Fanning the expert reads out: 0.296 → 0.400 tok/s, the biggest single round so far. D100 had shown the expert path's read half is latency-bound (1.08 GB/step at 0.83 GB/s, ~520 small preads issued one at a time), so DC-118 gives the provider a hint — preload(experts:shape:), called once per layer where the router's choices are known — and the adapter fans the misses across threads into the expert bank. The warm reads deliberately do not count as requests: the loop that follows is the requester and counts a hit. Alternated, one binary: 3.383 / 3.382 → 2.503 s/step, mix.read 1.671 → 0.761, trace digest identical in all four arms — and the bank without the fan-out is 4.109 s/step, so the win is the concurrency and the bank is only its staging area. The default is therefore 512 MB (not D98's 0, which was right for a bank that had to earn its keep by hitting). Instrument caveat, recorded because it would mislead otherwise: with concurrent reads the per-thread counters sum past wall time — read 3.867 s and unpack 5.206 s inside a 2.503 s step — so they measure thread-time under overlap rather than elapsed time; the reader's counters are write-locked for the same reason D101, DC-118, ExpertProvider.swift, MixtureOfExperts.swift
2026-09-18 mix.read is two costs and neither is the disk; and the operator authorised taking code from the sister project. SourceTiming had existed since the reader was written — its own comment says the caller needs to know whether those seconds are the device or the unpacking — and no caller had asked. Now the CLI writes it: on an 8-step 35 B-A3B decode, read 1.298 s/step (38.8%) and unpack 1.353 s/step (40.5%), 1.08 GB/step off the device (0.8 GB/s — not the limit) against 6.5 GB of fp32 materialised per step and discarded. That settles the next target: the unpack is what DC-120 exists to remove. The same round measured the GPU matmul and the disk watchdog stopped it at 3.94 GB free with swap at 1.71 GB — the head's 1.017 GB of bf16 becomes 2.03 GB of fp32 and the kernel wants an equal MTLBuffer, ~4 GB on a node with ~4.5 GB usable, so SHARD_GPU_MATMUL=1 is unmeasurable here until the fp32 array is gone. And the operator now allows taking code from TinyTitan, which is Apache-2.0: AGENTS.md and THIRD_PARTY_NOTICES.md record that attribution is the price — a root NOTICE (Copyright (c) 2026 André Borchert and the turbo-fieldfare notice), the licence text, marks on modified files — and that tools/check_provenance.py, which today refuses a third-party copyright line, is changed deliberately in the commit that first takes code (DC-123) D100, DC-113, DC-123, SourceTiming, datacenter-generate
2026-09-18 Threading the element-wise passes: 0.242 → 0.290 tok/s, digest unchanged — and a matmul attempt reverted for moving bits. decodeRaw is the LM head's bf16 conversion and the dense path's norms: a map, where every element is a pure function of its own bytes, so splitting it cannot change a value. It is now threaded through one knob shared with D94's unpack (DecodeThreads, SHARD_DECODE_THREADS=1 for the A/B). Alternated, one binary, four runs: 4.119 / 4.162 → 3.448 / 3.465 s/step, load 1.097 → 0.356, with the trace digest identical in all four (89d654ff54b0fd03). The matmul was attempted in the same round and reverted: parallelising (row, column block) looks free — each output is its own dot product — but the version moved bits and the repository's own reused-buffer comparison caught it, so Ops.orderedMatmulVectorized is the original body verbatim and DC-122 carries the failure. A speed-up that cannot be justified numerically is not one. The top phase is now mix.read at 1.71-1.77 s (50%), which is where the next work goes D99, DC-122, DecodeThreads.swift, Install.swift, ThreadedOpsTests.swift
2026-09-18 The expert bank was built, measured, and it answered the question the other way round. D97 read our hit rate of 0 as a lifetime problem: the bank was constructed inside loadLayer and dropped with the layer, so cross-token reuse was impossible. That was true and is now fixed — ExpertBank holds slices for the whole generation, byte-budgeted, LRU, per-projection cap and peak, with hits, misses, elements read and resident peak reported into metrics.json for the first time. Then the measurement said capacity is not the lever either. Five alternated 24-step runs on the real 35 B-A3B: 0.0% hits at 0, 512 and 1024 MB with identical 29,192.4M elements read, and the bank on slower in both pairs (4.020/4.132 s/step off against 4.139/4.236 on; 4.335 at 1 GB). The arithmetic closes it: 773 requests per step and 12.5 MB per slice, so one token's working set is 387 slices per projection — 4.83 GB, ~9.66 GB for both — and a 537 MB bank has a reuse distance nine times its capacity. A cache that could hit here would need 10-20 GB, which is not an 8 GB node. So the default is 0 (as D89 did for the layer cache), and the levers become the ones that do not need capacity: hide the reads (DC-121, the reference's prefetch ring) and cut their latency (DC-118). D31's "hit rate 0 at every size" was right about the workload as well as the lifetime; the lifetime fix is what made the two distinguishable. An inference worth stating as one: at 1-3 GB on this class of model a pure cache gives the same ~0%, so the reference's 60-85% hit column most likely counts prefetch-ring adoptions rather than cache hits — which is exactly the mechanism worth porting D98, DC-119, ExpertBank.swift, ExpertBankTests.swift
2026-09-17 The sister project's decode architecture, read and mapped. The operator authorised reading TinyTitan for approach, so the gap to 7 tok/s was studied at the source rather than guessed at. Its decode is a GPU-streaming design: PreadExpertStreamer streams experts with pread and fans misses across threads; ExpertResidencyTable is a GPU-visible (slot, state, generation) table so the kernel finds a resident expert in GPU memory; ExpertPrefetchRing stages predictions into raw-byte Metal slots and adopts them only if the router selects them; and affordableExpertCacheBudget clamps any requested cache to physicalMemory / 2 with the launcher's 30% warning on top. Ours has none of the four — and its expert bank is built inside loadLayer and dropped with the layer, which is why the measured hit rate is 0: the lifetime, not the size, and the reference's own curve (60.5% → 79.8% hits, +37% from 1 to 3 GB) says what that is worth. Four tasks recorded in order (DC-118 parallel misses, DC-119 generation-scoped bank plus the missing expert metrics, DC-120 int4 GEMV from resident slots, DC-121 prefetch ring), with the 4 GB row of the operator's table as the memory warning docs/reference-tinytitan-decode.md, D97, DC-117..DC-121
2026-09-17 The operator re-scoped the work: 7 tok/s on one node before any more network measurement — and the reference curve says why the gap is ~30x. The engine is at 0.230 tok/s on a Mac mini M2; the sister project measures a 35 B-A3B at 4-bit at 5.164 / 6.019 / 7.075 tok/s for 1 / 2 / 3 GB of expert cache on a comparable 8 GB M-series node, and 2.756 at 4 GB, where a wired cache plus dense weights, KV and prompt cache stop fitting in 8 GiB and the machine swaps against it (host wait 55 ms → 219 ms per token, GPU occupancy 42.8% → 17.4%, and 3.167 → 2.756 → 2.429 degrading run over run). Its verdict is bandwidth-bound, the lever is expert I/O, and its "at most 30% of physical RAM" warning is calibrated by that data. Reading our own code against it produced the finding that matters: our expert slot bank is built inside loadLayer and dropped with the layer, so on a cached decode it cannot hit across tokens and its hit rate is structurally 0 — D31's measurement was right and the cause is the lifetime, not the size. That, the switched-off GPU (33-46% busy over there), the complete absence of async I/O (0% hidden against 8-21%), the k-strided single-core matmul (~1/16 of usable bandwidth) and a head that materialises 2.03 GB of fp32 per token are the five gaps, in the order they cost time D97, DC-117, AGENTS.md
2026-09-17 Dense models are outside the design — measured, not assumed. The question was an overview across Qwen3.5 4B, 9B and Qwen3.6 35B-A3B, so the 4 B's real checkpoint was downloaded to a helper MacBook M3 (it is not a node), quantized there (426 tensors, 4,204,789,760 weights, 3.34 GB at 6.36 bits/weight), and the install was copied here to measure. It is slower than the bigger model: decode 5.870 s/step = 0.170 tok/s against the 35 B MoE's 0.230 tok/s on the same node — because the MoE activates ~3 B of its 35 B, while a dense model re-reads its whole payload every token: 18,638,208,000 bytes in four steps, 0 hits, since the default 1 GiB SHARD_DENSE_CACHE_MB cannot hold 3.34 GB. And it cannot be distributed at all: the shard plan divides experts, and a dense model has none. The 9 B is worse — 5.5 GB of payload against ~4.5 GB usable RAM. Tooling gaps found on the way: tools/quantize.py needs safetensors and torch directly, and a torch-free helper fails on the first shard until both are installed. The 4 B artifacts are deleted — the checkpoint and install on the helper (12 GB recovered) and the copied install here — so the evidence is the recorded measurement in D96, not a path that no longer exists D96, DC-116, README.md
2026-09-17 First release: v1.0.0, arm64 binaries for macOS 26 or newer. The engine now carries an identity a build can enforce — VERSION is the authority, sources/DatacenterEngine/Version.swift is generated from it, every tool answers --version, and both tools/version.py --check (a gate) and Package.swift refuse a disagreement. The second one was measured rather than assumed: a mangled mirror slipped past an incremental swift build because SwiftPM caches the compiled manifest, so the caveat is written down and the gate is the always-on half. tools/release.py is the packaging step RELEASE.md asked for: dry run by default, gates first, a clean scratch build with the warning scan, lipo -archs asserted on the binaries extracted from the archive rather than the build directory, one checksum beside the archive, and a --publish that refuses notes which do not quote the digest it just computed — a refusal with eight tests, because a check that has never failed is not yet a check. The release measures 0.23 tok/s prefill and 0.23 tok/s decode on one Mac mini M2, and 0.41 tok/s decode across four — 1.7x, bit-identical to a single node CHANGELOG.md, tools/release.py, tools/version.py, VERSION, RELEASE.md, docs/repository-decisions.md (D95), DC-115
2026-09-17 load was running on one core of eight — and the 35 B now runs at 1.70-1.74x over a single node. D88 measured load at 30.5% of a cached step and D93 at 47.9% of a cluster step, and D89 showed caching the decoded values loses to memory. What none of those said is the obvious thing: the engine is single-threaded (64-98% CPU on an eight-core machine) and load is ~1 G parameters of int4 constants dequantised with SIMD4 on one core — the same constants on every token and every node. The change is one loop: rows of an int4 tensor are independent, so the row loop in InstallFile.dequantizeInt4 is now spread across the cores, which changes which thread computes a value and not how it is computed (D63's rule). SHARD_DECODE_THREADS=1 restores the single-threaded path so the two are compared on one binary (D62's lesson). Bit-exactness was already guarded by the strongest check available: Int4UnpackTests compares the vector path against dequantizeInt4Scalar bit-for-bit over more than fifty shapes. Single node, alternated: load 1.33 → 0.56 s/step (2.4x), step 5.14 → 4.36 s, with head and mix.read untouched. Cluster, four alternated runs, every one bit-identical on all four nodes, loads 1.6-5.5: 1.74x and 1.70x with eight threads, 1.66x with one — so the objective is met, and the shard work alone clears 1.5x. Load hurts the ratio (four nodes see a spike, the baseline is one), so these are lower bounds. What it does not say: the gate's own roadmap target of ≥3x on a quiet farm is not asserted — every figure is an observation_only run with its loads beside it. And it retires a plan: DC-113's GPU GEMV was written down as the way to remove load; the CPU path took the same phase down with no device, no protocol and none of D91's silent-CPU-fallback risk docs/repository-decisions.md (D94), Install.swift, Int4UnpackTests.swift, DC-053, DC-113
2026-09-17 The head is vocabulary-parallel, and the cluster went from 1.13x to 1.36x. The ratio is (replicated + experts) / (replicated + experts/n + exchange), so the lever is the replicated work — and the head was the largest piece of it at 1.05 s/step on every node. Row v of the head is a dot product over the hidden width and depends on no other row, so each node now computes only its vocabulary slice, applied inside the existing headBlockRows loop so its rows are computed by exactly the same multiplications in the same order. The split is derived from the node count rather than declared in the plan, and VocabSlice is asserted to be a partition — every row owned exactly once, for vocabularies 0 to 248,320 over one to five nodes — because a row nobody computed stays zero and is neither an error nor a difference. Then the slices are gathered, so every node ends with the identical full logits array: the argmax, the margin and the trace digest are untouched, which matters because the gate compares tokens and digest. The gather is chunked and that is not an optimisation: a full slice is about a megabyte, and if every node sent one before reading it, the sends would fill the socket buffers and every node would block in send waiting for a peer blocked in send — so the vocabulary is cut into fixed windows, the same windows everywhere, each node sending one 16 KB frame per window and an empty frame where its slice misses, which is what keeps the reads aligned. Measured, four nodes, bit-identity intact: step 4.790 → 4.021 s, head 1.045 → 0.26-0.32 s, ratio 1.13 → 1.36x, and all four nodes now identical to the millisecond. The gap to 1.5x is 0.36 s/step — against a load phase that is 47.9% of the step. One number is recorded rather than used: the ledger's exchange_seconds (1.4-1.9 s/step) exceeds the ff phase containing it (0.575 s) even though the marks are right and the phases sum to the step, and D90/D92 both reasoned from that counter docs/repository-decisions.md (D93), VocabSlice.swift, HeadSliceWire.swift, ShardedForward.swift, ModelCache.swift
2026-09-17 The exchange is 99.6% waiting — and the waiting is the farm, not the plan or the protocol. D90 blamed the blocking, in-peer-order receive in allReduce, and acting on that would have been guessing: a blocking receive and a peer with nothing to send look identical from outside. So the timer was split — encode, send, receive, merge — and the segments cover 100.0% of the total: on a four-node run encode 0.002 s, send 0.007 s, receive 1.11-2.31 s, merge 0.002 s. The wire format and the protocol cost 9 ms per step; 90 reduces at 13-26 ms each is one node waiting for its peers to have something to send. Then the hypothesis was tested and falsified: four interleaved runs of contiguous against round-robin (the engine already supported interleaving; the gate now exposes --distribution with the values validated and four tests pinning the owner lists) gave 1.06, 0.96, 0.86 and 0.88× — indistinguishable — with a 24% run-to-run spread (slowest node 5.135 → 6.359 s/step), and the run with the slowest cluster step is the one whose peer sat at load 18.7. On this farm the cluster step measures the farm; no distribution fixes a peer doing something else. The same day's quieter run measured 1.13×. So a 1.5× figure cannot be certified without a quiet window — which is what the gate's --quiet-load 1.0 rule is for (D38), and why every figure carries its loads. What is in our hands: shrink the replicated work (the ratio is (replicated + experts) / (replicated + experts/n + exchange), and head at 1.05 s/step is identical on every node) and overlap the wait with prefetching that is not on the critical path docs/repository-decisions.md (D92), ShardExchange.swift, run_m2_gate.py, run_m3_gate.py, DC-051
2026-09-17 The cluster's cost is the exchange, and the exchange blocks on peers one at a time. The gate now forwards SHARD_PROFILE/SHARD_LAYER_CACHE_MB to every node (forwarded_environment, one rule and three launch sites, four tests) — without it a profiled run would have profiled the local node and called the other three unprofiled, or handed one node a different cache budget and called the difference a cluster result. Then the per-node profile of a real four-node run: the plan is demonstrably working — reads fall from 6.281 GB to 2.31-2.40 GB per node, mix.read from 1.80 to ~0.49 s/step, mix.gateup from 0.25 to 0.06 — and the exchange eats all of it: 4.250, 4.136, 0.672 and 3.739 s/step on nodes 0-3, that is 71.5%, 69.5%, 11.2% and 63.1% of the step. Four nodes doing identical work, one waiting 0.67 s and three waiting ~4 s. The cause is in the code: ShardExchange.allReduce sends its frame to every peer and then receives sequentially, in peer order, blocking — a slow peer does not delay itself, it delays everyone queued behind it, and it is paid forty times per generated token (~105 ms per layer, which is not a LAN round trip but a LAN round trip behind a busy node). The fix is contained and touches no arithmetic (frames merge by key in a deterministic order, so when one arrives cannot change a result): receive concurrently, and skip peers that own none of the chosen experts — roughly six of eight are remote, so one or two peers matter per layer, not three. That is the 1.5x, with the head shard (1.05 s/step, identical on every node) as the margin docs/repository-decisions.md (D90), ShardExchange.swift, run_m2_gate.py, run_m3_gate.py, DC-051
2026-09-17 The decoded-layer cache works, and on this node it loses — measured, so the default is off. D88 left load (30.5% of a cached step) as the target, and the shape is the first design question: not an LRU, because a decode sweeps every layer in order once per token, so the layer evicted is always the one about to be asked for. LayerWeightCache holds layers across the steps of a generation, counts the bytes from the arrays it holds rather than from a formula, returns a layer too large for the remaining budget without keeping it, and writes budget/bytes/layers/hits/misses into metrics.json per node — DC-052's done-when. It is bit-identical, asserted tensor by tensor. Then the measurement: 0, 256 MB, 1 GB and 2 GB alternated on one binary (conditions in sequence drift together, D62) gave 5.327/5.348, 5.322/5.328, 5.459 and 5.595 s/step — 5.46 and 5.60 s at 1 GB and 2 GB — with head +0.20 s/step and attn.core +0.27 on a machine already swapping at 2.1-2.4 GB. The cache did exactly what it was built to do (48 hits = 16 held layers × the three steps after the first, load falling as the budget rises, about 0.035 s per layer per step — 1.4 s of a 5.42 s step across all 40) and still lost to the resource it spends. So the default budget is 0 and SHARD_LAYER_CACHE_MB is the knob for a node with headroom: the D31/D62 shape again, a mechanism that saves the work it aimed at and can still lose. For the 1.5x target this is not the lever — the shardable and reducible parts are mix.read 33.0%, head 19.2% (vocabulary-parallel), mix.gateup/down 6.6%, and the exchange at 0.89 s of a cluster step docs/repository-decisions.md (D89), LayerWeightCache.swift, ModelCache.swift, DC-052
2026-09-17 The cached step's own breakdown — and it says the 3x target is not a sharding problem at all. The profiler is now wired through the path M3 actually measures: decodeOne takes one and marks the sequence path's phase names, generateCached runs a fresh profiler per step (so a phase's seconds are that step's, not a running total with the gap between steps folded in) and adds the reports, Generation carries the result, and datacenter-generate writes profile_seconds — with profiling: as a parameter so the tests turn it on without the environment variable. On the real install, the exact baseline command: 21.795 s over 4 steps = 5.449 s/step, phases summing to the wall, mix.read 33.0%, load 30.5%, head 19.2%, attn.core 9.6%, mix.gateup 4.6%. The three big phases are not the same kind of cost. The dense payload is read once for the whole run (1.0437 GB held, 4888 cache hits), so load is not I/O — it is the same ~1 G parameters of constants dequantised and released on every token, 160 times per generation, on the CPU while the GPU idles; head is cached weights and a matmul; and only mix.read is device-bound (~0.33 GB/step in 1.80 s, 184 MB/s). Since a plan divides mix.read alone, perfect, instantaneous, cost-free sharding is worth 1.09x on the measured step — the same conclusion as the forward's 1.39x, from the step that matters. So the attack is named: a decoded-weight budget, which is DC-052's done-when, not more sharding docs/repository-decisions.md (D88), ModelCache.swift, Generation.swift, DC-052, DC-051
2026-09-17 The first phase breakdown, and it settles the scaling question. The profiler was already built — nine phases in Qwen3_5Forward and ten inside MixtureOfExperts — but nothing the throughput gate runs reported it, so ProfileMetrics.fields() now single-sources the fields (and makes profiler off mean no fields at all, not zeroes: nil is not zero). Turning it on for a release trace on the real install gave the same digest (b0d382dbabf36df0…) and, over 40 layers and 17.363 s: mix.read 6.521 s (37.6%), attn.core 4.102 (23.6%), head 2.339 (13.5%, the LM head's 1.02 GB), load 2.338 (13.5%), mix.gateup+mix.down 1.805 (10.4%), everything else ~0.26 s. Two thirds of a forward is reading — mix.read + load + head = 64.5% — which is DC-107's I/O-bound finding at phase resolution. And it makes the target arithmetic rather than opinion: a four-node expert plan divides only mix.read, so perfect, free sharding of it is worth 6.521 × 3/4 = 4.891 s — 17.363 → 12.472 s, 1.39x, before any exchange cost. A 3x target cannot be reached by sharding experts on this design, and the measured 0.93x is what remains of that 1.39x on a busy farm. The lever is the replicated reads (dense 13.5% + head 13.5%), the attention core at 23.6%, and the read path itself — mix.read moves ~33 MB per layer in 163 ms, about 200 MB/s, the same range as the whole-node figure. What it is not: a five-token full-sequence forward, not the cached decode M3 measures — the cached path has no marks of its own yet (D86, D87) docs/repository-decisions.md (D86, D87), ProfileMetrics, DC-107, DC-051
2026-09-17 M3's first real number, and the bug the run found first. The operator authorised a throughput run on the busy farm, so the first thing run was the command run_m3_gate.py's own docstring shows a reader — and it failed on every peer, after copying the 21.7 GB install to each. stage_remote copies the install under its own name (m1-install) while both gates launched the peer with the literal ./install: one name spelled in two places, which only matches when the local install happens to be called install. --remote-install worked all along and every earlier cluster run used it — the flag that worked hid the path that did not, and the path that did not is the one a reader copies. stage_remote's docstring had drifted with it, claiming the install "is not copied here" while its own code copies it. The fix is one shared rule, remote_install_path(), used at all three call sites, with four tests — three on the rule and one that reads both gates' sources and refuses the literal \"./install\" anywhere, because the defect was the two halves disagreeing and a test of either half alone would have passed. Then the measurement, once the path was fixed: bit-identity on all four nodes, and 0.93x (baseline 5.662 s/step against the slowest node's 6.086 s/step). The gate labels it itself — observation_only, busy_farm, the loads recorded (node3 at 9.31), "this is not a gate result and the threshold was not asserted". The per-node metrics explain it, and this is the budget DC-051 has been waiting for: a node reads 2.65x less payload per step (0.592 GB against 1.570 GB, because the dense 0.261 GB is replicated and the experts really are 3.95x fewer) and is nevertheless slower; the exchange is 0.893 s/step (15.0%) moving 4.49 MB at 5.0 MB/s effective — 1.63 ms per term, latency-bound, not bandwidth-bound. So >=3x is not what an expert plan delivers on this design: the step is not expert-read-bound, which is DC-107's I/O-bound finding with numbers on it. What remains unmeasured is a per-phase breakdown of the ~3.8 s/step that is neither payload read across the plan nor exchange docs/repository-decisions.md (D83, D84), tools/run_m2_gate.py, tools/run_m3_gate.py, tools/test_run_m2_gate.py, DC-051, DC-053
2026-09-17 The Swift CI job built and tested, but nothing in CI compared that to the documented counts. The unexamined surface this round was .github/workflows/, and both files are in good shape — the standard-library job runs run_all_gates.py --skip-swift`` rather than a list maintained beside the tools, clones the wiki to check its tables, and reports a check it cannot run as not checked rather than as a pass. One hole was real: swift.yml ran `swift build` and `swift test --no-parallel` and never compared the result to the numbers the documentation claims, and the other workflow cannot, because it skips the Swift half and the claims gate then honestly reports it not checked. So "200 tests, 0 skipped, 0 failures" was compared to a real run nowhere in CI — a test deleted from the suite would have left every document correct about a suite that no longer existed, with both jobs green. The Swift job now runs the whole gate set, which is what its own header already said it was ("the engine's gates, on a clean machine") and which it was not doing. Stated plainly: CI still fails today by design, because `macos-26` carries Xcode 26.x below the manifest's floor and the toolchain gate stops the job — the change is fidelity for the day an image ships Xcode 27, not a green build now. A distinction worth keeping: a workflow's commands do not need the documented-command gate (`D81`), because they run — a renamed flag fails the job loudly — while a document's commands never run, which is why they were worth gating. Same words, different instrument, because a different thing happens to them docs/repository-decisions.md (D82), .github/workflows/swift.yml
2026-09-17 The last copyable surface is checked now: documented commands against the tools that take them. The repository's trap — check the form a reader copies, not only the prose — had been widened three times (the claims gate's own flags and paths in D67, the gates' invocations in D76), and what was left is the surface a reader meets first and copies most often: the commands in docs/, README.md, AGENTS.md, RELEASE.md and the wiki. The audit found nothing wrong — 45 flags, all accepted, and no command naming a script that does not exist — and two things about reaching that number: a first, narrower pass found only 32, because it skipped the wiki and did not join backslash continuations, so the instrument's own reach decided the answer again; and a silent surface looks identical whether it was checked or not, which is why the count is worth publishing. So it became a gate, check_documented_commands.py, holding three properties that decide whether it is honest: a tool that cannot be asked is NOT CHECKED, not passed (install_reader.py exits non-zero for --help, and the gate says so rather than counting its flags as fine — the one outcome it must never produce is silence dressed as success); prose that names a tool is not a command, or the pattern would invent flags to check; and a wrapped command is still a command. The same pass answered a second question: there is no dead code — all fifty non-test tools are referenced outside their own tests, and the twelve files that appeared unreferenced are all test_*.py, which unittest discover finds by pattern rather than by name, so nothing was removed because nothing should be. This is the first widening whose audit came back clean, and both outcomes are recorded because both are results docs/repository-decisions.md (D81), tools/check_documented_commands.py, tools/test_check_documented_commands.py, tools/run_all_gates.py
2026-09-17 DC-085 re-audited: the two install formats are not interchangeable, and the claim now has file-level evidence. DC-085 is the one open row whose action is tracking rather than timing or missing weights, so its audit was re-run. The checkout cannot answer the "releases" half — it is a source drop with no git history, dated before the original audit — so what can be audited is the snapshot, and what is worth re-checking is the claim AGENTS.md makes about it. Re-verified: its own reference states its layout in the first line of the function that consumes it — its own qwen35_reference.py, dequantize(): "bits-wide unsigned lanes packed low-first inside each uint32, one BF16 scale and bias per group" — returning grouped * scales + biases, with dequantizeInt4Affine on the Swift side; ours reads each code through _signed() (two's complement in four bits) and computes (code - zero) * scale with an int8 zero and an fp32 scale; and the containers differ (GTurboFormatV1, GTurboExpertV1, GTurboLayerV1, GTurboManifest*V1 against our install.json plus data.bin). So the conclusion is arithmetic rather than opinion, and there is no defect to report upstream, because its reader and converter agree with each other — the finding is that the formats are different, which is what makes the two projects independent implementations rather than copies. The evidence is now in the statement rather than only in a decision record: THIRD_PARTY_NOTICES.md recorded the code review and not what the format claim rests on, and a position a reader cannot check has to be taken on trust. One boundary, stated rather than glossed: this is a recorded audit, not an automated check — the artifact lives in another repository — while our own side is automated and the provenance gate inspects 158 files with no attribution to account for docs/repository-decisions.md (D80), THIRD_PARTY_NOTICES.md, DC-085
2026-09-17 I6's strongest form had never been run: the install's payload digests were an assertion the tooling did not honour. verify_install.py --digests hashes every payload against the manifest's sha256, and it had never been pointed at the real install — the manifest has carried those digests (and slab_sha256, whose whole point is checking one slab instead of a whole tensor) since the format was written, and nothing had ever compared them to the bytes. Measured, then decided: the whole install is 693 payload digests, 21,700,655,616 bytes, 26.69 s, 29,392,896 bytes peak RSS, zero problems, with free disk steady at 8 GB and swap unmoved, because it is one sequential pread pass and never a page-cached mapping — the envelope the tool's own docstring describes, now confirmed by a run rather than by its author's intent. And then it became routine, which is the part that matters: the milestone check now runs it, with the digest pass included rather than left to a separate invocation, --sample N to bound a quick run, and the verifier's output inherited rather than captured so the evidence appears in the run instead of behind a summary line. The check has teeth and they are tested: a copied install with one byte flipped is reported by the digest rather than by the geometry — the difference between a payload being the right shape and being the right bytes. What was verified for the first time is not a claim about the engine or the quantisation, but that the artifact every M1, M2 and M3 claim is computed from is the artifact the manifest describes, byte for byte docs/repository-decisions.md (D79), tools/check_milestones.py, tools/verify_install.py
2026-09-17 The install verifier had never been run by a gate — and the first time it was, it was wrong. I4 and I6 both rest on tools/verify_install.py, which only ever ran by hand. The milestone check now verifies the install before trusting a digest computed from it — a trace that reproduces its recorded digest on an install whose policy coverage or provenance is broken is a green light over a broken artifact — covering schema, roles, policy, tiling and provenance, and deliberately not the payload digests, which are a full read. Then the new step failed, and it was right to: the verifier reported the checked-in fixtures as mis-tiled. The rule was stricter than the format — it demanded exact contiguity, while D13 records the payload as 64-byte aligned (InstallWriter.add pads, because that is what SIMD4<Float> loads and a pread want) — so every install with a tensor whose size is not a multiple of 64 was called broken while the engine and the contract read it fine. It looked right because it had only ever been pointed at the real install, whose tensors are all far larger than the alignment and so tile exactly: D69 and D73 again in a different tool. The fix gives the alignment one home (install_reader.ALIGNMENT, used by the writer that pads and the verifier that checks) and the rule became no overlaps, and no gap as large as the alignment — plus a second correction the tests produced: a tensor's start need not be aligned, since the writer guarantees that but a reader does not need it. Two tests encoded the stricter rule and were updated, not deleted (their gaps are now a whole alignment wide, preserving what each is for), and a new test pins the other side. A verifier that cries wolf about a valid artifact teaches its reader to ignore it — and both installs now verify cleanly: the fixture at 55 tensors, the real one at 693 tensors tiling 21,700,655,616 bytes, after which the milestone check reproduces M1's digest and exits 0 docs/repository-decisions.md (D78), tools/verify_install.py, tools/install_reader.py, tools/quantize.py, tools/check_milestones.py
2026-09-17 The wiki's landing page carried a FALSE status, and three narrownesses hid it — each masking the next. .wiki/Home.md said "M1 has run ... at exit 0 and two byte-identical re-runs, and its gate is still open" and quoted 105 Swift and 174 Python tests, while the suites report 200 and 384. It was not stale-but-hedged like the headers fixed in the previous round: it was wrong. Three separate narrownesses kept it invisible, and fixing any one alone would have left it exactly where it was. (1) It was not a claim document — CLAIM_DOCUMENTS was a hand-maintained list of six, and the trap this module already names ("a gate whose configuration is a list will go stale") had bitten a second time in the same file; the fix is the one the trap prescribes, discover the pages, with News.md excluded by rule because a log records what was true when it was written. (2) Its phrasing was not recognised — the house style is **N tests, M skipped**, the page wrote "105 Swift tests pass (2 skipped" — so even once read, neither count matched. (3) The claim was inside a blockquote and wrapped, so the marker > sat between two words of the same sentence and line-at-a-time reading could not see it. The widened instrument then found the false claim before it was corrected — claims 105 Swift test(s) in prose; the suite reports 200 — which is the evidence that widening it was substantive rather than cosmetic; the numbers were then made true, the status replaced with what rounds 56-61 established, and the gate now checks 148 claims against 126 before this round, twenty-two of them on three wiki pages nothing had ever read. Two tests pin the exact shape that hid the claim. Every one of the three is the same defect in a different costume — an instrument narrower than the claim — and the repository has now found it four times: the flags, the list, the phrasing, the markup. In every case the gate reported success, because a claim it cannot parse is a claim it does not disagree with docs/repository-decisions.md (D77), tools/check_status_claims.py, tools/test_check_status_claims.py, .wiki/Home.md
2026-09-17 Both status headers were stale in the same way — leading with a failure that rounds 56-60 had already resolved. AGENTS.md — the file every agent reads — said of M1 "but that pass does not hold today", over a divergence that D56 had explained and D73/D75 had since re-verified in full; README.md said "what remains open is the tok/s measurement and D12" while AGENTS.md records that D12 was decided from a measurement (D31: the slot bank is sized from a budget, because the measured hit rate is 0 at every size). Both now lead with what is true and evidenced: M1's gate passes in both of its forms on all five frozen prompts — engine against a contract reading the same checkpoint (b8c976c5e7ba8816…) and the same install (b0d382dbabf36df0…), every comparison 83 tensors, 0 differing elements, 40 discrete decisions, peaks 3.60-3.79 GB and 0.98-1.41 GB. The history is kept, reordered as history, because it is why the claims are shaped as they are — the int4 projections, D55's declared cost, DC-112. This is the file's own trap ("the public status of this repository has swung three times, and only the latest is evidenced"): a status claim that points at a superseded measurement is wrong in the direction that matters, and the fix is the same one the trap prescribes — point at a command and its output, and make the newest evidence the one a reader meets first AGENTS.md, README.md, docs/m1-gate.md (D73, D75)
2026-09-17 M0's gate had no test — and a flag that disappears cannot be seen by a fixture. M1, M2 and M3's gate scripts each carry an instrument test ("so the gate is never untested code"); M0's did not, which mattered when D74 changed the invocation M0's gate makes, adding --uncached, with nothing watching. The new test drives the same script on the 236 KB fixture, so both halves run: byte equality against the contract and the discrete decisions against the reference implementation through trace_capture.py — contract IDENTICAL, oracle discrete MATCH, worst relative 1.92e-07, with the test asserting the smallest margin is positive, because a margin of zero is an argmax that happened to land right rather than a decision that was made. The subtler half: a fixture is small enough that a mapped read changes no number, and --uncached is byte-identical to the mapped reader by construction — so no byte-identity assertion can notice the flag disappearing, which is how it went missing for a whole era. The invocation had to become evidence: both gates now define the flags once, use that list for the command, and record it in the report, and the tests assert it — M0 checkpoint --uncached, M1 checkpoint --stream-experts --uncached, M1 install --stream-experts (absent on purpose, because an install reads uncached by construction and the install branch is chosen first). After this the class is assertable rather than merely discoverable — a gate that stops passing a required flag fails a test naming the flag, where three of the last eight rounds found such a defect only by reading code against a document docs/repository-decisions.md (D76), tools/test_run_m0_gate.py, tools/run_m0_gate.py, tools/run_m1_gate.py
2026-09-17 DC-114 fixed — and M1's gate now passes against the INSTALL too, all five frozen prompts. M1's restated claim (D56) is the engine against a contract reading the same install, and the gate could not finish that comparison: a five-token prompt had not completed in twenty-five minutes. The cost was one loop, measured before it was touched — _dequantize_row visits every value in Python, 0.253 s for one real expert, and a forward fetches ~2,560 of them. The fix follows the line the modules already draw: install_reader.py is standard-library only on purpose, so the offsets stay there (int4_block_bytes, now shared with the scalar path, so the layout is computed once) and the arithmetic moves to numpy in install_source._dequantize_int4_block. Bit-identical, argued and then measured: a small integer times a float32 is exact in float64 and rounds once, so float32 rounds to the same value; a full expert, a padded 64x2048 and a 32x2048 matrix all match under view(np.uint32), and the expert drops 0.253 s -> 0.013 s. A test pins the agreement over both fixtures' 25+ int4 tensors, including the MoE expert stacks. Then the gate ran: capital 17.40 s, arithmetic 54.51 s, code 67.63 s, repeat 86.90 s, long 98.45 s — all IDENTICAL, 83 tensors, 0 elements, 40 discrete decisions, GATE PASSED, digests b0d382dbabf36df0… on both sides, peaks 0.98-1.41 GB. Both forms of M1's gate now pass on this node. The install declaration moved 1.5 -> 2.0 GB because the measured peak is 1.41 GB and 6% is a coincidence rather than a guard. DC-114 is closed. No speed claim is made — the operator deferred benchmarking, and these seconds are what it takes for a gate to finish docs/repository-decisions.md (D75), tools/install_source.py, tools/install_reader.py, tools/test_install_source.py, DC-053
2026-09-17 The gate document's open question is answered by measurement: there was no engine drift. docs/m1-gate.md worried that "the drift is on the engine's install path", because a fresh engine trace printed b0d382dbabf36df0… where the stored pair printed b8c976c5e7ba8816…. Round 56's full run settles it: point the engine at the same input the stored pair used — the checkpoint — and it prints b8c976c5e7ba8816… again, byte for byte, after the GPU unpack became the default and every contract matmul went through a chooser. The two digests are two inputs, and D55 measures the distance between them. The gate doc now records the whole thing: all five frozen prompts engine-against-contract on the real checkpoint (capital 34.94 s / arithmetic 89.65 s / code 115.92 s / repeat 128.13 s / long 146.26 s, peaks 3.60-3.79 GB, GATE PASSED, 83 tensors and 40 discrete decisions on every one); that the gate passed neither of the flags this same file calls the survivable pair until D69 and D73, so its checkpoint runs mapped 67 GB; that the contract alone is 66.20 s and 0.397 GB once fixed; and that the declarations differ on purpose — 4.2 GB for a checkpoint because the engine's trace peaks at 3.60-3.79 GB, 1.5 GB for an install against a contract measured at 1.21 GB. A paragraph that said the remaining step "needs a machine with more memory" is corrected: it has been taken here, and what remains machine-specific is the install path's speed (DC-114). The invariants audit's I1 and I3 now cite the five-prompt evidence, and tools/milestones.json records the checkpoint form beside the install claim it deliberately uses docs/m1-gate.md, docs/invariants-audit.md, tools/milestones.json
2026-09-17 The "which reader, with which flag" question now has one home — and the audit found M0's contract couldn't be asked at all. D69 added --stream-experts to an invocation that omitted it, D73 added --uncached to the same one: two rounds, one shape. So this round audited every tool that invokes a contract instead of waiting for a third instance. ordered_qwen35_trace.py — M0's contract — had no --uncached at all: line 51 was source = SafetensorsSource(args.snapshot) and nothing else, so every M0 run mapped its checkpoint and run_m0_gate.py could not have passed a flag that did not exist. The 2B checkpoint is smaller than the 35 B one, which is why it never produced an incident, but the rule has no size threshold. The fix is one home: tools/contract_source.py holds open_source(snapshot, uncached=…) — install, else pread, else mapped — and add_uncached_argument(parser); both trace CLIs call it, so the 2B contract gains the flag and the install branch, and M0's gate and check_engine_contract.py now pass it. --stream-experts is deliberately not shared, because it is a property of the family rather than the reader: qwen3_5 is dense and has no stream_experts parameter, so a flag there would do nothing while reading like a capability — and a test asserts it is declared exactly where experts exist, encoding the reason. Seven tests, the important one being byte-identity: the 2B contract CLI is run twice on the fixture, mapped and through pread, with every tensor's sha256 compared. 381 tests, OK. ordered_reference.py still maps and is named as an exclusion rather than quietly skipped: it reads the small 0.6B reference models, not the pinned large ones. A hidden coupling also went with it — test_ordered_qwen36_quant.py reached SafetensorsSource through the CLI module, and now calls the shared rule docs/repository-decisions.md (D74), tools/contract_source.py, tools/test_contract_source.py, tools/run_m0_gate.py
2026-09-17 M1's gate PASSED on the real 35 B checkpoint, on this node, all five frozen prompts — and the other half of the safe pair was missing. docs/m1-gate.md records that both flags together are what make a real-model contract run survivable ("swap flat at ~1.2 GB", against "swap to 5.1 GB and disk to 8.0 GB in a minute" without them). The gate passed neither on the checkpoint path: D69 added --stream-experts, and this round found the path still going through safe_open, which maps 67 GB — exactly what the node's rules call the old hazard, while uncached_safetensors.py sat there saying "never mmap". Measured before changed: the contract on the checkpoint is 66.20 s and 397,344,768 bytes peak RSS (0.397 GB), against the 4.16 GB on record before either flag. Then the gate ran: capital 5 tokens 34.94 s, arithmetic 33 89.65 s, code 40 115.92 s, repeat 60 128.13 s, long 67 146.26 s — all IDENTICAL, 83 tensors, 0 elements, 40 discrete decisions, GATE PASSED, peaks 3.60-3.79 GB, and the first prompt's digest b8c976c5e7ba8816… is the one recorded on 2026-09-16, still reproduced after the GPU unpack became the default and every matmul went through a chooser. Expert traffic 2,240-8,040 requests and 3.5-12.6 GB read, hit rate 0 on every prompt — D31 from another direction. A correction within the hour: I re-based this path's declaration on the contract's 0.397 GB and set 1.5 GB; the gate's own 3.79 GB peak proved that wrong the same day, because a checkpoint run's peak is the engine's trace, not the contract's. Back to 4.2 GB with the measurement that justifies it — an under-declared guard admits a job the machine cannot take. D65's "belongs on a bigger node" is superseded for the checkpoint path; the install path remains (DC-114) docs/repository-decisions.md (D73), tools/run_m1_gate.py, DC-053
2026-09-17 The class of D71, not just the instance — an index mapping now has a check. After fixing a formula, the question this round asked was where else that kind of defect could hide, and the answer was that the check which should have caught it did not exist. The manifest is provably consistent: for every tensor in the real install (693) and the fixture (36), nbytes equals prod(shape[:-1]) rows of padded_columns (or of shape.last when zero means not padded) — 0 mismatches, 20,695 MB accounted for against a known ~20 GB install. So the payload, geometry() and verify_install.py were all right, and the wrong number was in the one derived quantity nothing checked: _rows_per_index, the map from the checkpoint's index space onto the install's rows. The fix is an invariant rather than another patch. check_index_mapping asserts that the install's rows are the indices times the rows one index takes (shape[0] * factor == geometry().rows) and that one index contributes exactly its trailing dimensions (factor * shape.last == prod(shape[1:])) — for D71's tensor the first fails as 8 x 16 = 128 against 256 rows, the half-expert that cost two rounds. It runs for every tensor when an InstallSource is opened, so it fails at the point of use in every caller. Verified both ways, because a check that cannot fail is decoration: the real install and the tiny one open cleanly, and a test pins the D71 geometry and reinstates the wrong formula and asserts the refusal, naming both numbers. 374 tests, OK. Two rounds went into bisecting what twenty lines of arithmetic catch at open — the generalisation is that a quantity derived from a checked one needs its own check, because the check on the input says nothing about the derivation docs/repository-decisions.md (D72), tools/install_source.py, tools/test_install_source.py
2026-09-17 DC-113 fixed — the bisection worked, and it showed my own previous fix was wrong. Quantising one role at a time on the tiny install isolates it exactly: every role is IDENTICAL except the two expert stacks. The decisive number was in the manifest — the quantiser records nbytes: 9472 for a (8, 32, 32) stack padded to 64, which is 256 rows of 32, the layout Install.swift computes — so the Python reader was wrong. _rows_per_index divided the flattened trailing width by the padded row (16) where the truth is the product of the middle axes (32), so it returned half of every expert; that is the cannot reshape array of size 512 into shape (1, 32, 32) of D69. And D69's fix was the wrong half of the bug: I changed columns when the row count was wrong, misled by _width's docstring claiming every trailing dimension is flattened — the payload's nbytes is the authority and it was one command away. The real model never showed it because shape.last and padded_columns are both 512 and the middle product is 2,048, so both formulas agree: every real geometry is provably unchanged, and the table that proved it found a second latent bug in the same function — the convolution was read as factor 4, columns 0, when it is one row per index and four values. Verified by the count moving: the install-mode gate case that was DC-113's expected failure now passes, and the suite reports expected failures 2 where it reported 3. DC-113 is closed docs/repository-decisions.md (D71), tools/install_source.py, tools/test_run_m1_gate.py
2026-09-17 DC-113 is int4-specific — and my performance hypothesis died on measurement. DC-113 runs in 0.04 s, so it can be interrogated: both sides on the tiny checkpoint are IDENTICAL (7 tensors, 0 elements, 2 discrete decisions), and both on the tiny install differ first at layer.00.hidden_out by 18,676,984 ULP while embed.out and layer.00.hidden_in match. A difference that large with a matching input is structural. Layer 0 is a Gated DeltaNet layer, and its int4 tensors are exactly the padded ones — in_proj_qkv/in_proj_z/out_proj 32 wide stored at 64, the expert stacks 16 and 32 wide stored at 64, group size 64 — so every one is a partial group on a padded row, the DC-087 family, in a geometry the real install does not have and therefore never tested. Next step is a bisection, not a guess: quantise one role at a time and compare the two readers' values. And the other half of the round killed my own theory. M1's gate could not finish a five-token prompt in 25 minutes, so I suspected the recorded list[float] trap; measured on one real expert, iterating rows with no accumulation (0.320 s), the list[float] (0.315 s) and a preallocated fp32 array (0.335 s) are the same, and the rate saturates at 17.7 MB/s — so the cost is neither allocation nor syscall overhead but the per-row crossing into Python, 2,048 per expert, against the Swift path's ~1 GB/s (DC-114, a timing-phase item). What that means is a cost, not a verdict: M1's gate is functionally able to run against an install since D69, and needs hours per prompt here until that read path is fixed docs/repository-decisions.md (D70), DC-113, DC-114
2026-09-17 M1's gate can run against an install — which it never could — and the tiny fixture found a latent shape bug. D65 said the gate needs a bigger node because it hands the checkpoint to both sides; that was true of the checkpoint form and never of the install form, which M1's claim was restated to in D56 and which needs no mapping. The gate could not do it because its contract command line never passed --stream-experts — while install_source.py's own docstring said the contract used it — so the contract died with InstallSourceError on the first expert stack. The gate now passes the flag and discovers checkpoint vs install from the directory (install.json's family / config.json's model_type) rather than demanding another flag. The flag is correctness for the install, which refuses to materialise a stack, and memory for the checkpoint: those sources do not refuse, they materialise — which is why this gate declares 4.2 GB. The declaration stays, because lowering a guard to match a measurement is how guards get weakened. Then the fixture earned its keep: with the flag the gate failed on the tiny install with ValueError: cannot reshape array of size 512 into shape (1, 32, 32), because _materialise took its row width from the original tensor's last axis instead of the install's own padded. For the real model those are equal (2,048 either way) and for the fixture 32 against 512 — correct for the 35 B model, wrong in general, and found by the new install-mode case driving a path no checkpoint case reaches. Two process notes, both mine: I released the heavy-job lock while the gate held it, and the next tool refused and printed the holder back at me; and Python buffers stdout when redirected, so a 12-minute run's log held only its first warning — long background runs use -u docs/repository-decisions.md (D69), tools/run_m1_gate.py, tools/install_source.py, tools/test_run_m1_gate.py, DC-053
2026-09-17 M3's four-node mesh re-verified — and the functional half is now runnable on its own. D66 left the mesh alone because a peer was at load 14; the farm went quiet (2.5 / 1.4 / 1.0) and this round ran it. Four machines, current binaries, a full mesh driven by the loaded plan, three peers reading their own installs so only the binary and plan crossed the network: every node IDENTICAL in tokens and digests — [11751, 11, 264, 3177], the same four D65 verified for the single node. The gate itself now expresses the operator's split: --functional-only runs the cluster and the bit-identity check and stops, computing, printing and recording no speedup, and nulling the timing fields rather than omitting them so the output cannot be mistaken for a measurement. It also skips the quiet-farm precondition and says that it skipped it, because a functional run makes no throughput claim. Two of my own defects, both caught by the run: the new early-return wrote its report into a directory the normal path creates later (a traceback after four IDENTICAL lines), and I hid that exit code with a pipe to tail — the gate crashed and my command reported exit 0. That is the third time this session the recorded | tail lesson has bitten, and the fix is shape rather than care: capture to a file, then read $? docs/repository-decisions.md (D68), tools/run_m3_gate.py, AGENTS.md, DC-053
2026-09-17 "The form a reader copies" now includes paths — and all fifty-three of them were already right. A previous round widened the claims gate to read the --swift-tests flag beside the sentences, because that is the number most likely to be pasted. Paths had the same argument and no check: the link gate validates [text](path), and a bare tools/run_m1_gate.py in a code block is not a link. The gate now requires every tools//docs//sources//tests/ path with an extension in the six claim documents to exist, with two rules that are properties rather than lists: a template is not a claim about a file (docs/release-notes-vX.Y.md in RELEASE.md is a shape, and every path here is lower case, so mixed case is skipped by rule rather than by exception), and an empty scan is not a pass — the checked count went from 73 to 126, so "no problems" means fifty-three paths were looked at. The result is negative and that is the good outcome: the README, AGENTS.md, the tracker, the roadmap, the testbed and the architecture pages are all consistent with the tree, which is now enforced instead of hoped for. The same audit caught one stale thing in the gate itself — its --help example still read --swift-tests 184 --python-tests 200, values nothing checked but a reader may copy — and those are placeholders now, because an example that cannot go stale is better than one that has to be maintained docs/repository-decisions.md (D67), tools/check_status_claims.py, tools/test_check_status_claims.py
2026-09-17 The cluster claims were not gate-checked either — I2 and M2 re-established on the current engine. check_milestones.py re-runs the single-node trace and nothing else, so I2, M2 and M3's functional half rested on manual runs made before the GPU unpack became the default and the matmul chooser landed. Re-run two-machine, real model, current binaries, with the peer reading its own install (a functional test, not a data migration): the single-node reference, node 1 on the peer over TCP and node 0 here all digest b0d382dbabf36df0…, and trace_diff reports IDENTICAL — 83 tensors, 0 differing elements, 40 discrete decisions for both nodes. A 256-expert plan over two contiguous halves, each node reading its own half (2.05 GB apiece), the reduction contract exercised 40 times per node over 797 and 803 terms. One observation recorded and not asserted: the all-reduce cost differed threefold between the nodes (1.50 s here against 0.48 s there) — the local node is the connector, it had just run the reference, and this is the 8 GB host. The four-node mesh deliberately did not run: a peer is at load 14, and that claim belongs to the quiet window with the throughput gate. Asking "what has not been re-checked?" is now two for two docs/repository-decisions.md (D66), docs/invariants-audit.md (I2), tools/run_m2_gate.py, M2's tracker row
2026-09-17 M1's "identical generated tokens" is verified — and now a gate, not a memory. The record has always said generation is 0.108 tok/s cached and 0.0374 uncached with identical generated tokens; the rate was measured and the equality was observed by hand, and the engine has changed since. Re-run on the install: both paths produce 11751,11,264,3177,34756,364,1141,8807 with identical top-2 margins at every step — including 0.1339 and 0.0445, narrow enough that a small numerical difference could move them — and the first token matches the argmax D40 measured independently. run_m1_gate.py now runs the cached path too and requires the token lists to be equal (I3: a discrete decision, no tolerance), and an unreadable line fails the gate rather than passing, because [] == [] is a true statement about nothing. Verified by the gate's own fixture test, which drives the whole script on the 236 KB model. What this node cannot do is run M1's gate itself: it hands the checkpoint to both sides, so it maps 70 GB — the mechanism behind two panics — and that run belongs on a machine with more memory, in the timing phase. Also: my first run lost the token line to tail -3. The recorded | tail lesson has a second form — a claim being verified has to be captured in full docs/repository-decisions.md (D65), tools/run_m1_gate.py, tools/test_run_m1_gate.py, DC-053
2026-09-17 The tiled kernel is bit-exact and still slower — and the profile says kernels are not M3's lever. D63 diagnosed the GPU matmul's slowness as the 8 KB stride between adjacent threads and specified a tiled kernel. It is written, it is right — a warp now reads 32 contiguous floats as one cache line, each thread still accumulates ascending in k, and the 288-shape grid plus the real trace (b0d382db…, trace_diff IDENTICAL) hold — and the diagnosis was wrong: coalescing changed nothing, and the tiled and naive kernels cost the same. What is left is the path around the kernel: an encoder, two copies and a synchronous wait per call, about 4 ms per dispatch across ~210 calls in mix.gateup. Two traps the grid caught rather than reasoning: a return before a barrier left the first element of every short row right and the rest wrong (surviving lanes read rows whose owners had exited), and loading each lane's own row is correct and is exactly the stride being removed. The strategic reading: mix.read + load + head are 12.73 s of an 18.75 s forward — 68% — at the ~1 GB/s sequential floor, so GPU compute is not on M3's critical path and the lever is sharding the reads, which M2 built and M3 demonstrated. The kernels stay opt-in as a tested asset and the work stops here; DC-033's Metal clause is recorded with its measurement and the task is closed. Also: the first tiled measurement showed every phase identical between on and off — because the flag's default had been inverted the round before, so both runs were the CPU. Check that the conditions differ before believing either docs/repository-decisions.md (D63, D64), MetalMatmul.swift, DC-033, DC-107
2026-09-17 The GPU matmul is bit-exact and slower — so it is opt-in, and round 45's head number was drift. Routing every contract matmul (38 call sites) through the chooser left correctness untouched — the real trace is b0d382db… on the GPU and on the CPU, trace_diff IDENTICAL, 200 tests — but the speed went the wrong way. The first comparison looked like machine drift because every phase was up, including mix.read, which no matmul touches; alternating the conditions on one build resolved it: attn.core 4.05 → 5.55 s, mix.gateup 1.17 → 1.62 s, head 1.74 → 1.86 s, with the untouched I/O phase identical. Reproducible to a tenth. The lesson is the method: conditions run in sequence drift together, which is exactly how the earlier head figure (CPU first, GPU second) came to show a gain that was not there. The cause is the access pattern: one thread per output walks w rows k*4 bytes apart, so a warp touches 32 cache lines to use 4 bytes of each, while the CPU's vectorised path walks k contiguously and reuses x. The fix is known and safe: a threadgroup-tiled kernel, coalesced into shared memory, which cannot change a result because tiling changes which thread accumulates and not the order. Until then the matmul is opt-in (SHARD_GPU_MATMUL=1) — a kernel that is slower is not a default. What still holds from round 45: the head reads 1.02 GB of its own weights, and that is its floor docs/repository-decisions.md (D62, D63), MetalMatmul.swift, DC-107
2026-09-17 The head's matmul on the GPU: bit-exact over 288 shapes, 26% off the phase — and the rest is I/O. With D61's arithmetic settled the kernel is simple: one thread per output, k ascending, acc + metal::fma(x, w, 0.0f), no reduction and nothing to reassociate. MetalMatmulTests asserts bit-identity against Ops.orderedMatmul over 288 shapes — chosen so output counts that are not multiples of four are in the grid, where the CPU's four-wide vector takes its tail — plus the head's real shape and cache reuse. The chooser lives in the forward rather than in Ops, because Ops is the contract and should not know about a device, with a million-multiply-add threshold and a fallback that cannot change a trace because the two are bit-identical. The real trace is b0d382db… with the head on the GPU and on the CPU, trace_diff IDENTICAL. Measured by phase: head 1.74 s → 1.28 s — while the totals said nothing (16.3 against 16.4 s, the wrong way round, because the reads vary by more than the saving; a phase profile resolves what a total cannot, which is D60's lesson from the other side). What remains is 1.02 GB of head weights at ~0.8 GB/s, the same sequential floor as the expert reads, so the head is documented as at its limit. attn.core at 4.05 s did not move and is next: small activations, arithmetic-dominated, the right case for this kernel docs/repository-decisions.md (D61, D62), sources/DatacenterEngine/MetalMatmul.swift, MetalMatmulTests.swift, DC-107
2026-09-17 D10's open question, measured: no math mode is safe for a bit-exact GPU matmul — fma(x, w, 0) is. The comment on the unpack's pipeline asked whether .relaxed was enough for the GEMM that follows, or whether the accumulation had to be guarded. Measured with a one-ULP discriminator rather than a shape: .fast, .relaxed and .safe all contract a * b + acc into one fma (.safe only keeps a product in its own statement apart, so safe/plain is still fused), which would put a bit-exactness claim one ULP out. accumulator + metal::fma(x, w, 0.0f) reproduces the contract under all three modes, so the GEMM can share .relaxed with the unpack. The trap is worth more than the answer: the same probe over a 257-term dot matched in all nine mode-and-formulation combinations, including the contracting ones — a test written against a long dot, the obvious way to test a matmul, would have declared the GPU exact and been wrong. Both bit patterns of the discriminator are asserted in the test so it cannot pass by being unable to tell the difference docs/repository-decisions.md (D61), tests/DatacenterEngineTests/MetalMathModeTests.swift, DC-107
2026-09-17 The GPU unpack is now the default, and a profile run corrected the priority list. D59 left it behind a flag on the argument that a default is a policy; that is sound for a change that could move a digest, and this one cannot — it is asserted bit-identical over the DC-087 grid, a real fixture tensor, the partial-row payload the row path assembles and buffer-cache reuse, and end to end the real trace is b0d382dbabf36df0… with the decoder either way. Verified after the switch: 193 Swift tests, 0 failures (their fixtures now run through the GPU path), default 16.1 s, SHARD_GPU_UNPACK=0 17.3 s, and trace_diff against the stored scalar trace IDENTICAL for both. A SHARD_PROFILE=1 run then corrected DC-107: mix.read is 7.23 s of a 19 s trace — the read 3.24 s plus the unpack 4.53 s — not the 3.45 s the record held, and the unpack was never at a limit; the read is (3.24 s for 3.06 GB is ~0.95 GB/s, the sequential floor). The remaining order is attn.core 4.05 s, load 2.22 s and head 1.76 s, and the head is the better first kernel: a plain GEMM, and the easiest to make bit-exact docs/repository-decisions.md (D59, D60), sources/DatacenterEngine/Install.swift, DC-107
2026-09-17 The GPU unpack now beats the CPU path with the same digest — the cost was the plumbing. D58 measured it at 38.2 s against 18.2 s and named why: the row path concatenated the three sections into a Data, unpack sliced it twice more, each section was copied into an [UInt8] before reaching a buffer, and a fresh 8 MB output buffer was allocated per expert fetch — 2,218 times in one trace. A one-slot buffer cache, held under a lock across the dispatch and its completion so reuse cannot race a running kernel, plus one copy per section instead of four, gives 16.3 s — faster than the CPU — with trace_diff against the scalar run reporting IDENTICAL — 83 tensors, 0 differing elements, 40 discrete decisions. DC-107's target is met: faster, with the same digest. The numbers are wall-clock observations of one trace each rather than a benchmark, the timing phase stays deferred, and the default stays opt-in because a default is a policy that the frozen prompt set should justify. The cache's own risk — a buffer grown for one shape leaking into the next — is covered by a test that unpacks small, large, small and large again docs/repository-decisions.md (D59), sources/DatacenterEngine/MetalUnpack.swift, MetalUnpackTests.swift, DC-107
2026-09-17 DC-087 fixed and the Metal unpack is wired in — bit-identical end to end, and honestly slower. The GPU unpack has existed since D10 and was called by nothing but its own tests, because it disagreed with the scalar path on a row whose final group is partly filled. That reproducer now passes, and the acceptance grid was widened from a list that skipped every awkward width to one that includes 17, 33, 65 and 129 — bit-identical on synthetic payloads, on a real fixture install tensor, and now on the partial-row payload the engine's own row path assembles (the case every other test missed). SHARD_GPU_UNPACK=1 selects it and the default cannot move, because every recorded digest came from the scalar path: the real 35 B trace is b0d382dbabf36df0… with the flag off and on, and trace_diff between the runs reports IDENTICAL — 83 tensors, 0 differing elements, 40 discrete decisions. The cost is measured and it is not the kernel: 18.2 s scalar against 38.2 s GPU, because the row path concatenates the three sections and unpack copies them again, allocating an MTLBuffer per fetch across 2,218 fetches. That plumbing is the next target (DC-107). And a lesson that cost the disk: the watchdogs had been killed while tidying up the round before, so this run — which grew swap to 1.8 GB — took free disk to 0 GB with nothing watching. It recovered on its own and one bad sample is not a trend, which is why the watchdog acts on three; but the watchdog now runs with the job, always docs/repository-decisions.md (D58), sources/DatacenterEngine/Install.swift, MetalUnpackTests.swift, DC-033, DC-107
2026-09-17 I6 closed: the artifact can name its source — and the last partly is gone. The M1 install predated the provenance mechanism (source.files empty, repo holding a commit hash, revision saying "local"), so the artifact every gate reads could not say which weights it came from. Both values were discoverable from the cache path rather than needing a human: models--Qwen--Qwen3.6-35B-A3B/snapshots/995ad96e… names the repo and the revision. tools/repair_install_provenance.py repaired the metadata in place under three rules — the payload is provably untouched (only source and passes may move, every other key compared canonical-JSON-equal), the repair is recorded as a provenance-repair pass, and a disagreement is refused rather than overwritten. Result: 26 source-shard digests, verify_install clean where it used to print three complaints, and the engine's trace unchanged (b0d382db…), so every digest recorded in docs/ and on the wiki still stands. Cost: 1m19s with the disk flat at 9 GB. Also decided, with numbers: DC-112 — the DeltaNet's and attention's projections should return to bf16, because the logits say the int4 pass is narrow (argmax unchanged on all five positions, but the smallest margin is 0.27, top-8 overlap 5–8 of 8, mean logit difference 0.30), and the routed experts keep int4 where the size win lives. Execution is blocked on ~25 GB of free disk docs/invariants-audit.md, docs/repository-decisions.md (D57), tools/repair_install_provenance.py, DC-112
2026-09-17 M1 verified: the engine reproduces the contract byte for byte — the hunt's answer was that the two sides were reading different weights. tools/install_source.py points the contract at the install through the same dequantiser the Swift reader mirrors, and trace_diff then reports IDENTICAL — 83 tensors, 0 differing elements, 40 discrete decisions, matching digests b0d382dbabf36df0…. Against the bf16 checkpoint contract the same engine differs by 40 discrete and 1 float, which is the declared, measured effect of int4 — D55: the DeltaNet's projections and attn.q/k/v/o are int4 in the install, worth 0.0156 max / 24.5% median relative at layer 0 — and is tracked in DC-112. The milestone's claim in tools/milestones.json is restated to the falsifiable form, so the re-check verifies the right thing; DC-111's hunt is closed and its row removed. The finding that cost the most: an install flattens trailing dimensions, so expert.stack_gate_up is 262,144 rows of 2,048 and one expert is 1,024 consecutive rows — D32's "one expert is one row" holds for the checkpoint, not the install — and a source that assumed it failed loudly instead of producing a plausible trace. Cost of the run: 1,600 expert fetches, 878 MB and 19.5 min CPU at the twenty-minute mark, disk flat. Also learned the hard way: running the gate battery beside that job made a Metal test fail, which is the "never a heavy job alongside an engine run" rule docs/m1-gate.md, docs/repository-decisions.md (D55, D56), tools/install_source.py, tools/milestones.json, DC-112
2026-09-17 M1's gate had two inputs, so it could not be tested — now it has one. D55 showed the divergence was the weights rather than the arithmetic, which means the gate's comparison was ambiguous by construction: any disagreement could be either. tools/install_source.py points the contract at the install, through the same dequantiser the Swift reader mirrors, so both sides see identical weights and a difference is a difference in arithmetic. The finding worth keeping is the row space: an install flattens a tensor's trailing dimensions into the row length, so the real expert.stack_gate_up is 262,144 rows of 2,048 and one expert is 1,024 consecutive install rows — D32's "one expert is one row" holds for the checkpoint, not the install. A source that assumed it returned a thousandth of an expert, and failed loudly on the reshape instead of producing a plausible trace. Install.row_range also lands (fetching expert 255 no longer dequantises 254 experts), with ten tests and 359 in total. The full-model run is in flight — 1,600 experts through a per-value Python dequantiser, measured at 878 MB and ~19.5 minutes of CPU with both guards live, so its result is recorded when it is measured and not before docs/repository-decisions.md (D56), tools/install_source.py, tools/install_reader.py, DC-112
2026-09-17 M1's divergence is the quantisation policy, not the engine — and the proof is byte-exact. Five rounds of bisection ended by recognising that the two sides were never running the same weights: the install holds the Gated DeltaNet's three projections (linear.in_qkv, in_z, out) and attn.q/k/v/o as int4, by tools/quant_policy.json, while the contract is generated from the checkpoint's bf16. Running the reference's layer 0 on the identical hidden_in with the install's own int4 values reproduces the engine's attention output byte for byte (max abs 0, identical=True); with the checkpoint's weights it differs by 0.0156 max / 24.5% median relative, compounding through forty layers into the forty flipped decisions. The engine is correct — through the convolution, the gates, the l2 norms, the chunked delta rule, the triangular solves, the gated norm and the projection — which is stronger evidence than the original gate ever gave, and it also means the pass recorded on 2026-09-16 cannot be reconciled with these artifacts. M1's claim and the policy are in conflict, which is a decision: DC-112, blocked on about 25 GB of free disk for a rebuild that must not run on this node. Also this round, and my fault: two of my own comparison scripts grew to gigabytes on this 8 GB machine — one wrapped a streaming iterator in list(), the other asked dequantize for a tensor as a Python list[float] — and the second tripped the disk watchdog at 3.42 GB free, the exact shape of the 2026-09-16 panic. Disk recovered, three stale disk watchdogs were found still running, and the memory budget is now enforced rather than advised: tools/memory_watchdog.py, a 4 GB real-memory ceiling, three readings before acting, a MEMORY_STOP marker, and seven tests docs/repository-decisions.md (D54, D55), tools/memory_watchdog.py, docs/m1-gate.md, DC-111, DC-112
2026-09-17 The divergence is inside the delta rule — and it is not a rounding difference. Two more capture points per side (delta_core before the gated norm, gated_norm_out after it; 223 tensors on both, and both agree that 30 of 40 layers are DeltaNet) localise it exactly: layer.00.hidden_in matches, layer.00.delta_core differs in all 20480 values — inside chunk_gated_delta_rule, with the input norm, the conv, the gates, the scale, l2norm and the repeat all eliminated. The character is what matters: median relative difference 4.3%, 90th percentile 29%, 2–4.7% on values above 1e-3 — where a 1-ULP difference is ~1e-7 relative. So this is structural, not a cast, a fused multiply-add or a Float-vs-np.float32 order, which eliminates the family of explanations the reading had left. And yet the body reads as identical, stage by stage, twice. The next comparison is the rule's inputs — query, key, value, beta, decay — because with a zero initial state and one chunk the core is intra_chunk_attn · solved and nothing else. A line-by-line read has now twice agreed with itself against a measurement GatedDeltaNet.swift, Qwen3_5Forward.swift, ordered_qwen35.py, D53, DC-111
2026-09-17 Reading the chunked rule to the end: every stage corresponds — so reading has hit its limit. Ten candidates are now eliminated: the convolution, the decay gate, the input norm, silu/sigmoid, the query scale (checked by number rather than by eye, because d**-0.5 rounded from float64 and 1/sqrt(d) in fp32 genuinely can differ in the last bit — they agree for every dimension this model uses, and a power of two is exact), l2norm, the triangular solve (whose own docstring says the order the rows are eliminated in is what the contract fixes; both ascend with one subtraction of a product per step), v_new/inter/the output, the state update, the grouped-query repeat (np.repeat and (head·repeats + i)·headK are the same interleaving), and the projection/conv order. No stage reads as different and the outputs differ systematically, so the cause is something a side-by-side read cannot see — a cast, a fused multiply-add, a Float-vs-np.float32 evaluation order, or an input verified only by reading. The next comparison has to be of numbers: capture inside chunk_gated_delta_rule and around the gated norm GatedDeltaNet.swift, ordered_gdn.py, D52, DC-111
2026-09-17 Reading the attention half: four stages out, and the two-path structure found. Rather than instrument deeper, I read the half D50 localised. Eliminated with reasons: the convolution; the decay gate (-exp(A_log)·softplus(a+dt_bias) on both sides); the input norm — the half's first op, and the one a systematic all-elements difference would most likely come from, so it was checked closely: both take 1/sqrt (the contract says in so many words that it is chosen over a hardware rsqrt because they differ in the last bits), both sum the squares ascending one addition per step, both apply the family's weight-offset norm (1+w)·(x·inv) — and the multiply order swaps, which is harmless because IEEE multiplication is commutative; and silu/sigmoid, where the contract's sigmoid branches at zero and both sides do, since a single-expression sigmoid would agree on positive inputs and differ in the last bits on negative ones. What the reading found instead: the contract names torch_chunk_gated_delta_rule:301, the engine has that as chunkedDeltaRule, and it also has a recurrent torch_recurrent_gated_delta_rule:440 path reached with rows: 1 — the decode case. A five-token trace takes the chunked path, so the divergence is inside it or the ops around it: intra-chunk attention with exp(-inf) masking, the cumulative-decay prefix sum, the per-chunk state decay, the state update, and the grouped-query repeat that DC-038 already had to fix in this path GatedDeltaNet.swift, ordered_qwen35.py, D51, DC-111
2026-09-17 M1's divergence is in layer 0's attention half — shown by an instrument, not an argument. Both sides now capture what is inside a layer, opt-in (SHARD_TRACE_INTERNALS=1 for the engine, --capture-internals for the reference), recording attn_out and ff_out per layer: 163 tensors instead of 83. Opt-in is a requirement, not a courtesy — the digest covers the tensor list — and the proof it is off by default is that the default trace is still 83 tensors with digest b0d382db…, byte for byte. One run then bought the elimination of the whole mixture path: layer.00.hidden_in identical, layer.00.attn_out differs in all 10240 values, layer.00.ff_out differs — so the router flip and all forty layers of discrete decisions are consequences of something in the input norm or the Gated DeltaNet. Eliminated by reading: the convolution, and the decay gate (-exp(A_log)·softplus(a+dt_bias) in both, negation outside the exponential). Remaining: the chunked delta rule, the projections and L2 norms, and the gated RMSNorm, where the engine's own comment records that its bf16 oracle rounds the normalised value before the weight multiply and an all-fp32 contract does not. Also caught: my first capture patch hit the non-streamed call site, so the run produced 83 tensors and looked like a successful capture of nothing docs/m1-gate.md, Qwen3_5Forward.swift, ordered_qwen36.py, D50, DC-111
2026-09-17 Localising M1's divergence: the convolution is out, and I had misread the differ. Two corrections and one elimination. The correction is mine: DIFFERENT — 40 discrete, 1 float means one float tensor, not one element — all 10240 values of layer.00.hidden_out differ, and the quoted value is only the first. So the difference is systematic inside layer 0, not a boundary rounding, and the earlier reason for suspecting the convolution is withdrawn. A cheap probe — recompute the reference's router on the engine's layer-0 output — disagreed with the reference's own record, which is a broken instrument rather than a finding: hidden_out is the layer's final output, while the router ran on the mid-layer value the trace does not capture. That confirms a bisect needs capture points inside the layer. Eliminated by reading instead of running: the convolution, where both implementations index position + tap − (kernel − 1), accumulate in the same order and zero-pad identically. Candidates left: the delta rule's first step, the gates' exponentiation, and the gated RMSNorm's ordering — which the engine's own comment flags as the place where its bf16 oracle and an all-fp32 contract differ docs/m1-gate.md, docs/repository-decisions.md, D49, DC-111
2026-09-17 M1 fails on today's artifacts, and the divergence is one element of one layer. Two blockers had to go before the question could even be asked, and they were different: the reference memory-mapped a 67 GB checkpoint (D47), and it materialised a whole layer's experts — 3.2 GB in fp32, which drove swap to 5.1 GB and disk from 11.8 to 8.0 GB in a minute when first attempted. Routing the experts by index (the fact DC-032 used on the engine side: the leading axis is the expert) makes the contract run finish in two minutes with swap flat at 1.2 GB, and it is proven invisible: the streamed contract is IDENTICAL to the stacked one on the real model. With that, the verdict: fresh contract against today's engine, 40 discrete decisions and 1 float differ — the same as against the stale contract, so it is real. layer.00.hidden_in matches while layer.00.hidden_out element 0 differs by 28% relative but only 0.0014 absolute — a boundary difference, not a structural one — and layer 0 is a Gated DeltaNet layer, whose causal convolution has its boundary exactly there. The router flips marginally and the decisions cascade through all forty layers: I3's failure mode arriving through one element docs/m1-gate.md, tools/ordered_qwen36.py, D47, D48, DC-111
2026-09-17 The reference can now read a checkpoint without mapping it — and the attempt that proved it found the real blocker. The M1 gate could not be re-established here because safe_open memory-maps a 67 GB shard on an 8 GB node, the mechanism behind both panics. tools/uncached_safetensors.py serves tensor and rows by pread through F_NOCACHE, never mapping, and is byte-identical to safe_open on a real 4 GB shard — a row slice of the 1074 MB expert tensor and a 1-D norm, with every dtype the checkpoint uses, plus a hand-built file carrying a NaN, both infinities and a signed zero. The reference takes it behind --uncached. Then the run was attempted rather than reasoned about, and the blocker turned out to be underneath: with the page cache out of the picture, swap grew 2048 → 5120 MB and free disk fell 11.8 → 8.0 GB in sixty seconds, because the reference's own working set is 2 GB of fp32 for a single expert stack. Stopped deliberately, recovered to 10 GB, no panic. DC-111 now names memory, with the numbers tools/uncached_safetensors.py, docs/m1-gate.md, D47, DC-111
2026-09-17 The milestones now re-check themselves, and the M1 confusion is explained. The drift round 29 found was found by hand — thirty seconds of work nothing in the repository was doing. tools/check_milestones.py makes it repeatable and keeps two claims apart that look alike: the engine is stable (it still produces its recorded digest — a single-node trace, twenty seconds) and the engine matches the reference (which needs the contract, and on the 35 B model that means the checkpoint this machine cannot safely read). Its first full run: M1 reproduces with a stale contract (DC-111), M2/M3 reproduce with a current contract because their comparison is between machines reading one install, and M0/M4/M5 are not checkable here with the reason stated. The M1 gate's own document turns out to be inconsistent — a later table lists b0d382dbabf36df0…, which is what the engine produces today, while the status section still quotes b8c976c5e7ba8816… — and nothing had re-checked which was true now. A declared divergence stays visible without failing (a permanently red tool is one people stop reading); an undeclared one fails tools/check_milestones.py, tools/milestones.json, D46, DC-111
2026-09-17 The M1 gate is not current, and the re-check that found it took thirty seconds. A fresh engine trace on today's install compared against the stored contract is DIFFERENT — 40 discrete decisions and 1 float, where the two stored traces were byte-identical when written (a7c77b63…). The reference is not the stale side: its only change since is a two-line headroom guard, and it reads the checkpoint, which has not moved. The engine reads an install that was rebuilt at 04:50, and the dequantiser changed with D34 at 23:36. Settling it needs a GB-scale streaming read of the 67 GB checkpoint — the operation that took this node's free disk from 17 GB to 2.96 GB and helped panic it — so DC-111 records the status change and the exact discriminating step instead of risking the machine. M2 and M3 are unaffected: all four nodes read the same install and produce b0d382db…, and the recorded baseline counts still match to the byte. The lesson is general: a stored trace is evidence only while both sides that produced it are unchanged docs/m1-gate.md, docs/repository-decisions.md, DC-111, D45
2026-09-16 The memory floor and the one-job lock are enforced now, not just written down. The project's worst incident was a machine panicking — a 35 B engine at ~4.5 GB resident beside a 20 GB install build on a node with ~4.5 GB usable, thirteen swapfiles, and the hardware watchdog. Disk had a floor; memory and concurrency were rules in a document. tools/heavy_job.py makes them real, and the split is the design: a job whose declared peak exceeds usable memory is refused with the arithmetic (physical less the measured reserve), a machine that already looks short on reclaimable memory produces a loud warning but not a refusal (available memory is fuzzy, and the disk watchdog already made that mistake once), and being alone is a lock with stale detection — a claim whose process is gone is taken over with a notice, because a lock a crash leaves behind is one the next person deletes. Wired into quantize.py, run_m1_gate.py and run_m3_gate.py, declaring their measured peaks. Also DC-110 closed: the transient failure did not recur in ten consecutive runs on a busy farm — recorded as unexplained rather than fixed tools/heavy_job.py, D44, DC-110
2026-09-16 The failure rules were tested the way a cluster actually fails: a node was killed mid-exchange. The survivor failed fast and by name — connection reset by peer: the peer's process is gone, not merely quiet — which is last round's transport fix holding under a real kill rather than a simulated one. The run also found three things: a killed peer sends RST, which the transport was reporting as a generic socket error (it now has .reset, mapped from ECONNRESET on read and EPIPE/ECONNRESET on write, with the retry rule unchanged); the harness printed node 1 failed: followed by a blank line, because a killed process has no stderr; and .close() — added so a test could force RST — turned out to make any later use raise an Objective-C exception that Swift cannot catch, so a closed transport now refuses with a Swift error instead of crashing. Two of my tests were wrong and are rewritten or gone: asserting which of reset or closed a killed peer produces asserts the kernel's choice (the same test passed filtered and failed full), and the write-side twin is a race, not a contract sources/DatacenterEngine/ContributionTransport.swift, tools/run_m3_gate.py, D19, D43
2026-09-16 Two more tasks close on evidence that already existed, and the wire protocol becomes a contract document. DC-009 asked for the all-reduce wire protocol — framing, dtype, timeouts, retries, node-failure semantics — and its acceptance, a node that stops answering fails the run by rule and a retry cannot change the bits, is met by D17's canonical (token, expert) order plus D19's rule table and the tests that hold each row, now including the real-socket ones that found the mid-frame hang. Rather than point at a decision log, docs/wire-protocol.md is the ADR: the frame field by field, the two different geometry checks, the refusals that happen on the declared numbers before anything is allocated, and a table naming the test behind every failure rule. DC-080 — one policy file drives more than one family — turns out to be verified already: the policy test walks every committed fixture install, two families' worth, and asserts the union of their roles matches the policy in both directions, so a missing role and an unused entry each fail. 16 open tasks docs/wire-protocol.md, tools/quant_policy.json, DC-009, DC-080
2026-09-16 A documented failure rule did not hold on a real socket, and the run hung instead of failing. D19 says a peer that stops part-way through a frame fails the run because the stream is desynchronised — and ShardExchangeTests asserts it, with a peer that is a Swift object. Three new tests use a real socket, and one hung: readExactly polled until readable and then called FileHandle.read(upToCount:), which is readDataOfLength: and blocks until it has every byte. After poll correctly reported three bytes it waited for the other sixty-one, so a peer that stops mid-frame produced an infinite hang, no retry count, and no documented failure. read(2) returns what is there — which is what the poll was for — and the test that hung now finishes in 0.061 s against its 60 ms deadline. Re-verified after the change: 189 Swift tests, 0 skipped, 0 failures, and bit-identity on four real machines with the fixed transport, every node producing the baseline tokens [11751, 11, 264, 3177]. The mesh run that looked like a ten-minute hang was scp-ing a 20 GB install because --remote-install had not been passed; the harness now announces what it will copy, and the pipeline that swallowed its output is the lesson recorded one round earlier sources/DatacenterEngine/ContributionTransport.swift, tests/DatacenterEngineTests/TCPTransportTests.swift, D19, D42
2026-09-16 DC-036 met: the Swift gate runs on a clean machine, and the whole battery is one command. tools/run_all_gates.py runs every gate and — the part a list of commands in a shell cannot do — reads the test counts from the suites' own output and hands them to the claims gate, so the stale-count mistake stops being possible instead of being watched for. It earned that immediately: on its first run on another machine it reported 288 Python tests where the documentation said 276, because this round had just added twelve. Staged on a farm node that has never built this checkout, the battery passed — toolchain OK, tables OK, links OK, provenance OK, python 288, **swift 186 with 0 skipped and 0 failures**, status claims OK — which is the task's criterion: the farm carries Xcode 27 and Swift 6.4, so the Swift gate runs on a clean machine, and GitHub's hosted runner is precisely the thing that cannot. Two bugs the tool found in itself, both the same kind: it parsed the first Executed … line (a suite's six tests) and reported 13 of 186 — XCTest prints the total last — and its per-file Python run used a dotted module path that broke the sibling imports nine first-party test files rely on tools/run_all_gates.py, .github/workflows/markdown-links.yml, DC-036
2026-09-16 DC-084 closed: the recorded baselines are data now, and a fresh run reproduces every one of their counts exactly. tools/baselines.json + tools/check_baselines.py split them in two: a count (expert requests, payload bytes, cache hits) is deterministic, so it is asserted exactly and can be re-checked even on a busy farm; an observed value (seconds, memory) is reported with its delta and asserted only when a caller asks, on a quiet farm. The demonstration ran a trace and a cached generation at the recorded conditions and matched 2218 / 0.0 / 3,060,562,432 / 1,043,708,416 / 4277 to the byte, re-producing digest b0d382dbabf36df0… unasked — while throughput came out 0.1761 tok/s against the recorded 0.108 (+63.1%) on a quieter farm, which is precisely why observations are not asserted. Two bugs came out of running it: the applicability field was prose, so the two-node fixture's exchange_reduces was compared against a single-node run and a correct 0 was reported as a failure (fixed with a machine-checkable requires map), and derive crashed on the real step_seconds list because the unit tests had invented a scalar shape no run writes — the test now uses the recorded shape, and an unexpected one is NOT CHECKED rather than a traceback tools/baselines.json, tools/check_baselines.py, DC-084
2026-09-16 The invariants are audited, and the audit found a real gap in an artifact. I1–I6 had been assessed before M1 and M2 existed, so docs/invariants-audit.md now records where each one stands with its evidence and its falsifier: I1, I2, I3 and I4 verified (three milestone gates; two machines and four in a mesh; separate discrete-decision assertions; the quant policy and the shard plan as data), I5 not applicable yet (the checkpoint is bf16, so there is no vendor quantisation to transcode — it becomes live at M4, which has no weights here), and I6 partly. I6 is the find: the code records null rather than a placeholder revision and digests the source files, but the install on disk predates both fixes — revision: "local", files: {}, and a commit hash sitting in repo — and nothing had ever checked the artifact. verify_install.py now reports those three gaps on every run, the packer warns when its two provenance inputs look swapped, and check_status_claims.py requires every invariant to keep a section, a status and named evidence, so the audit cannot decay. The stale assessments in the brief response are marked, not rewritten docs/invariants-audit.md, tools/verify_install.py, tools/quantize.py, I6
2026-09-16 DC-018 closed: every node holds the same verified install — and the audit that came with it corrected a claim of ours. node1 (91 GB free) and node2 (70 GB) each took the 21.7 GB install at 115.8 MB/s — 3m07s each — and each was verified on the node itself: 693 tensors tiling 21,700,655,616 bytes, policy matched, 5 payloads (2.97 GB) hashed. Using the verifier away from this checkout found a real defect in it: the manifest records the policy's absolute path on the building machine, so a copied tool died looking for a stranger's home directory. It now prefers the file beside the script, says which file to copy when it cannot find one, and has a test for it. Two harness lessons too — my ssh pipeline (| tail) reported exit 0 for a remote crash, and the first fix failed because the policy had not travelled with the tool. Separately, auditing the sister project's quantiser settled that its int4 is unsigned-with-bias in GTurbo*V1 containers while ours is signed codes in its own container — each internally consistent, so no upstream defect, but this repository's documents had claimed it "holds the install format" and now say what is true (D39) tools/verify_install.py, docs/repository-decisions.md, DC-018, DC-085
2026-09-16 M3's gate is built, and it refuses to run on a busy farm — which is today. run_m3_gate.py asks every node for its one-minute load average before measuring anything and stops if any is above the threshold, naming what it saw: node4 3.00, node1 5.58, node2 2.44, node3 2.25. node1 at 5.58 is the gate detecting my own install staging on that node — the instrument working. Built is not passed, and the record says so: --allow-busy-farm records an observation that cannot pass the gate, with the threshold unasserted and the exit status decided only by bit-identity. Three rules are built in: correctness (tokens and digest) is asserted before speed and never traded for it; the ratio uses the slowest node, because a cluster step ends when its slowest member does; and a node nobody can ask is NOT CHECKED, not quiet. Its numbers come from each node's metrics.json (D37), so the join and staging are not counted as compute. The gate needs the install everywhere, so it is being staged on the two nodes that lacked it (node1 91 GB free, node2 70 GB, against 21.7 GB — neither near the 5 GB floor) tools/run_m3_gate.py, docs/repository-decisions.md, DC-053, DC-018
2026-09-16 DC-081 closed: the cluster's cost is accounted for, per node. ExchangeLedger counts what every all-reduce did — reduces, terms sent and received, bytes both ways, and the wall seconds inside the exchange — and it counts every attempt, because a retry that cost a round trip is part of what the run cost. Each node writes its own metrics.json and the harness prints a table; the local two-node run on the fixture shows 7/5 and 5/7 terms from the two nodes, so the table is self-checking — an all-reduce that dropped or duplicated a peer's terms would show as an asymmetry before any trace comparison. A sharded generation across two machines reports the same shape per node (10 reduces, 10 terms each way, the dense payload read once and served from cache 120 times) with tokens and digests still identical to the single-node reference, so the accounting changed no arithmetic. A node with no metrics is reported NOT REPORTED, not as zeroes. The seconds are raw observations on a shared farm, not a throughput claim: DC-053 owns that and needs a quiet farm sources/DatacenterEngine/ShardExchange.swift, tools/run_m2_gate.py, docs/repository-decisions.md, DC-081
2026-09-16 DC-013 closed: nothing here is copied, so no NOTICE transfers — and now that is measured rather than assured. The sister project is Apache-2.0 and carries a NOTICE (naming turbo-fieldfare); this repository is MIT. A read-only clone (19 MB, into .build/, never committed) was compared file by file: its modules are all TinyTitan* with no shared module name, four files share a basename and none shares more than 5 % of its non-comment lines, and the identifiers in common are format and model vocabulary (hiddenSize, vocabSize, safetensors, dequantize) — several of which come from the model's own config.json. The stronger evidence is positive: the install reader exists twice here, in Swift and Python, written independently and checked against each other, with its layout learned from the artifact (D29). What is shared carries nothing — a format is an interface, conventions are ideas, and the Apache-2.0 weights are never redistributed. Apache-2.0 §4(d) attaches a NOTICE only when the distribution includes the material. THIRD_PARTY_NOTICES.md records the position and the three conditions that would change it, and tools/check_provenance.py guards it offline: the file must exist, keep naming the relationships, and no source file may acquire a third-party copyright header. Python tests 211 → 219 THIRD_PARTY_NOTICES.md, tools/check_provenance.py, docs/repository-decisions.md, DC-013
2026-09-16 New gate: the documentation's claims are checked, and its first run found one — a true one. tools/check_status_claims.py compares the test counts each current-state document claims against the suites' actual output, checks that every Dnn cited anywhere in the repository has a definition in a decision record, and checks the release state; it reports how many claims it found and refuses to pass on zero, because a regex that stopped matching would make it a decoration. Its first run flagged News claiming 105 tests, 2 skipped — which is not drift but history: the entry was true the day it was written. So the gate now separates claims from records, and the false positive is a test. The CI job feeds it the tool suite's own count, parsed from the run that just happened, so the number cannot be a literal someone forgot to update. Four rows left the tracker with it: DC-004 (the CI gates), DC-034, DC-035 and DC-044 were all done and still open — work finished but unrecorded is the same defect as a claim recorded but wrong. Python tests 200 → 211 tools/check_status_claims.py, docs/repository-decisions.md, DC-004
2026-09-16 The GPU and CPU dequantisers are bit-identical, and the residue was three defects, not one — the Metal tests are no longer skipped. DC-087 had recorded the difference as "the sign of zero" with the denormal half settled by D11. Working it out found: (1) the GPU had never implemented D11 at all — it uploaded scales raw while the CPU and the Python reference read denormals as zero, and 0.722705 % of the real install's scales are denormal; (2) the sign of zero is emergent on both sides, so neither was the reference — the contract now says a computed zero carries no sign; (3) a NaN had no definition — numpy propagated a payload where Metal canonicalised — so a NaN is the canonical quiet NaN. The rule also had to stop being the additive idiom: value + 0.0f is correct IEEE and the Metal compiler folded it away, leaving nine values of 195 as -0.0. And the first fix went into one of two CPU implementations; the other kept the old behaviour for two shapes out of sixty. Evidence that nothing about the model changed: the same real-install prompt still writes 83 tensors, 40 discrete, digest b0d382dbabf36df0… — the digest the M2 gate recorded, byte for byte. swift test --no-parallel now runs 184 tests with 0 skipped sources/DatacenterEngine/MetalUnpack.swift, sources/DatacenterEngine/Install.swift, tools/quantize.py, docs/m1-decisions.md, DC-087, DC-090
2026-09-16 The repeat-read audit corrected its own author. D32 ended with a guess written as a finding — that the cache-off run's 2.78 GB per step against a 1.04 GB distinct set meant tensors were read twice inside one forward. So the guess got an instrument: a per-tensor count of whole-tensor payload requests. It says 611 tensors asked for more than once, layer 0's ×8 — and 8 is exactly the number of forwards (five prompt-replay steps plus three decode steps), so it is one read per forward. The arithmetic closes to the byte: 8 × 1,043,708,416 = 8,349,667,328, the cache-off total, and 1,043,708,416 is also everything the cache holds. No unexplained traffic, and the prompt's two passes are ModelCache's documented choice — the outputs come from the verified sequence path the gate compares, not from a second implementation of it. D33 records the correction, the instrument stays as observability (DC-081) sources/DatacenterEngine/Install.swift, tests/DatacenterEngineTests/InstallCacheTests.swift, docs/m1-decisions.md
2026-09-16 DC-106 closed: the dense backbone stays resident — 8.35 GB of reads became 1.04 GB, with output unchanged. The brief's premise was that the backbone stays resident while the experts stream; the engine re-read it per token (2.95 s of a 19.8 s forward). Measured from the install first: the payload is 21.70 GB — 18.87 GB of stacked expert banks and 2.83 GB of everything else, of which the embedding and the LM head are 1.02 GB each and are read a row at a time, leaving a per-layer backbone of just 0.796 GB. The cache holds the packed bytes, because 0.74 GB of that is int4 and its decoded form would need ~6 GB against a ~4.5 GB budget — DC-091 a second time. It sits on InstallFile.payload(...), the only whole-tensor path, which the stacked expert banks never take, so 18 GB of experts cannot fill it and a test asserts that. The A/B, same install and prompt, three cached-decode steps: 1,043,708,416 B read with the cache on against 8,349,667,328 B off (8×, 4,277 reads served from 1.04 GB resident) — with identical tokens and an identical trace digest (620f5acc777a0d6b…), because residency changes what is read, never what is computed. One surprise: with the cache off the run read 2.78 GB per step against a 1.04 GB distinct set, so the same tensors were being read more than once inside a single forward. No timings claimed sources/DatacenterEngine/Install.swift, tests/DatacenterEngineTests/InstallCacheTests.swift, docs/m1-decisions.md, DC-106
2026-09-16 D12 decided by measurement, twice over: the expert slot bank is now sized from a budget and the answer is one slot. The question had been argued twice — the literal 16 (8.05 GB across 40 layers against ~4.5 GB usable, an expected failure in SlotBudgetTests) and a cross-token bank that audit reverted for the same reason — so the missing thing was a number, and the number is a count: hits and bytes. On the real 35 B install, one five-token prompt, five bank sizes, one process each: 2218 expert requests, 0 hits, 0.0000 hit rate, and byte-for-byte the same 3,060,562,432 bytes read at 1, 2, 4, 8 and 16 slots. A bigger bank saves exactly nothing, because a bank lives as long as its layer and inside one layer call the router's picks do not repeat. So the capacity now comes from SHARD_EXPERT_BANK_MB (default 512 MB → one slot per layer) and the layer's own shape, with SlotBudgetTests asserting the constraint, the budget property, the refusal of an absurd budget, and the answer itself — and the expected failure deleted deliberately, as its own note asked. The measured win is elsewhere: DC-106, the dense backbone's residency, at 2.95 s of a 19.8 s forward. Hits and bytes are counts, not timings; the runs were functional and sequential with swap and free disk flat sources/DatacenterEngine/Qwen3_5Forward.swift, tests/DatacenterEngineTests/SlotBudgetTests.swift, docs/m2-decisions.md, DC-091, DC-092
2026-09-16 DC-083 closed: the wire audited against hostile input, and two real gaps found. The frame format was already defensive, but "already defensive" is a claim, so it was tested as one. First find: a frame can declare a geometry the receiver does not have — decode validated a frame against itself — and the terms reached OrderedReduction.accumulate, whose width invariant was a precondition. A mismatched peer crashed the node, from the network. Now: decode reports what the sender declared, ShardExchange checks it and refuses with geometryMismatch naming both, and accumulate throws instead of preconditing — belt and braces, the second of which matters on its own. Second find: the term-count bound was per-number, not per-product, so reserveCapacity could ask for terabytes; the reservation is now bounded by the bytes actually present. The audit is a seeded fuzz suite — 2,000 random buffers, 2,000 single-byte mutations, every truncation, hostile declarations, and a frame from another geometry — each with a two-second bound, because a duration bound is what can see an allocation. Posture recorded rather than assumed: one address, one port, only declared peers, no authentication and no encryption by design for a switched private segment — and the Architecture page says what would have to change first. Swift tests 167 → 176 sources/DatacenterEngine/ContributionWire.swift, tests/DatacenterEngineTests/WireFuzzTests.swift, docs/m2-decisions.md, DC-083
2026-09-16 DC-108 closed: the install has a reader in the other language, and all of it verifies. The gap this project recorded against itself was that the engine reads the install while the contract reads the checkpoint, so nothing in Python could ask whether the artifact the engine actually reads is what the packer claimed. tools/install_reader.py (standard library, streaming, uncached) and tools/verify_install.py now answer that: the 693 tensors tile the 21,700,655,616-byte payload exactly — so no tensor can be reading another's bytes — every role's quantisation is the one tools/quant_policy.json requires, and all 693 payloads hash to the manifest's digests: 26.3 s for 21.70 GB, uncached at 826 MB/s, free disk flat at 15 GB and swap unchanged either side. Two things worth reading: the format carries padded_columns = 0 for the 221 unpadded tensors (a reader that demanded a padded width first refused every norm), and the quantisation against the source checkpoint is still not checked, because the 67 GB source is not in this checkout — reported as not checked rather than implied by a green run. Python tests 179 → 200 tools/install_reader.py, tools/verify_install.py, docs/m1-decisions.md, DC-108
2026-09-16 DC-109 done, and reading first caught a defect that would have hidden behind a benchmark. ModelCache.decodeOne called the mixture directly, so a sharded --cached run would have computed only its own experts, all-reduced nothing, and produced a perfectly reasonable token sequence — the failure mode this project exists to eliminate, in the one path a throughput gate would have used. Fixed the D23 way: one entry point (Qwen3_5Forward.mixtureOutput), used by the sequence path and the cached path, so no call site can forget the reduce. datacenter-generate gained --plan/--config/--node and mesh joining moved into ClusterJoin, shared with datacenter-node. Evidence, real 35 B install, two machines, cached decode: both nodes generated [11751, 11], the reference's tokens, with identical trace digests (ec3fd8f114e88dc2…) and trace_diff IDENTICAL on both. No timings claimed — the wall clocks are in the log and the farm is shared. Two test bugs found loudly on the way: a hand-built socket triangle deadlocked (the test now joins through the production mesh rule) and free ports were discovered by holding the listeners that then could not bind sources/DatacenterEngine/ModelCache.swift, sources/DatacenterEngine/ClusterJoin.swift, sources/DatacenterGenerate/main.swift, tools/run_sharded_generation.py, docs/m2-decisions.md
2026-09-16 The whole farm runs one forward: four machines, one bit-identical trace. --mesh gives every node its own process on its own machine, joined in a full mesh rather than a star — because the all-reduce is pairwise (D17), a leaf in a star would hold only its own terms and the coordinator's, and isComplete would refuse the run. Joining is deterministic: node i connects to lower ids and accepts from higher ones, so no pair connects twice. All four nodes handshake, reduce and write 7 tensors, 2 discrete decisions, digest 58518422914cfe2b…, and trace_diff says IDENTICAL on every one. That is M3's functional half; its ≥3× throughput gate is not claimed — the farm is shared, nothing was isolated, and timings wait for the timing phase. Runs now use the internal LAN addresses rather than the node names, which resolve over the VPN at ~2.5× the round trip; the Testbed inventory carries them so the next run does not rediscover them tools/run_m2_gate.py, sources/DatacenterNode/main.swift, docs/m2-decisions.md, DC-053
2026-09-16 M2's gate passes on the real 35 B model, across two machines. A 256-expert plan over two contiguous halves, node 1 on node3 and node 0 on the development host, trace_diff on every trace: 83 tensors, 0 differing elements, 40 discrete decisions, matching digests b0d382dbabf36df0… — and that digest is the single-node M1 baseline's, produced from half the experts on each machine. I2 stops being an argument and becomes a measurement. The trace carries the logits, so identical logits make greedy generation identical by construction — a stronger claim than comparing generated text; the missing piece is a sharded datacenter-generate (DC-109), which the tok/s gates will need. No timings are claimed: the wall clocks are how long the functional test took, the farm is shared and nothing was isolated. The install was staged once (21.7 GB into a node's Downloads folder) because the development host already had it. One correction: an earlier decision record printed a node's address, and this repository is public — the line is scrubbed and the slip recorded; it survives in that commit's history, and removing it is a force-push the operator decides on. The tracker's M2 row and phase table now read Done tools/run_m2_gate.py, docs/m2-decisions.md, DC-045, DC-109
2026-09-16 M2's gate passes on two machines. tools/run_m2_gate.py --remote node1@node1 stages a node binary and the fixture on node1, runs node 0 here, and the project's own differ reports IDENTICAL on both nodes against the single-node reference — bring-up handshake, one all-reduce per mixture layer, digests 58518422914cfe2b… on both machines. The binary is built here and copied, so a peer needs no checkout and no build; ad-hoc signing survives scp, which is what makes one arm64 binary enough to turn a machine into a cluster node. It is the fixture, deliberately: the 35 B model needs its 20 GB install staged on a peer, the farm's nodes are shared with other work, and the standing rule for this phase is functional tests rather than benchmarks — so the same command with --install becomes the real gate when the timing phase opens, ideally over Ethernet rather than the VPN the names resolve to. Three harness bugs fell out of assuming the local case: two listeners on two hosts waiting for each other, a trace copied without -r, and output labelled by process index so a passing run read as a failing one. The Testbed now records that the nodes are shared, so cluster runs stay functional until the farm is quiet tools/run_m2_gate.py, docs/m2-decisions.md, DC-018, DC-045
2026-09-16 M2's gate runs as two processes over TCP and trace_diff says IDENTICAL on both. datacenter-node is one node as its own process — its own install reader, address space and socket — and tools/run_m2_gate.py starts two, has them handshake, runs two sharded forwards, and judges every trace with the harness M0 built: 7 tensors, 0 differing elements, 2 discrete decisions, matching digests 58518422914cfe2b…. Two threads in one address space share a heap; this does not. Along the way the harness's hardcoded family guess (tiny-qwen36 against the fixture's real qwen3_5_moe) was refused by the node, which is D20 working — but the refusal arrived as a Swift SIGTRAP, because an uncaught top-level try is a trap rather than a message; both sides fixed. The node's metrics are measured rather than derived: the elementsRead * 2 estimate removed in the first round of this work does not come back wearing a different name. Honest limits: the CLI is two-node, and the run is the fixture — the real model at two nodes does not fit on this 8 GB host, which is exactly what DC-045 is waiting for. DC-041 closed sources/DatacenterNode/main.swift, tools/run_m2_gate.py, docs/m2-decisions.md
2026-09-16 The engine runs sharded, and the trace is byte-identical to the single-node one. D23: a Qwen3_5Forward with a ShardExecution is one node of a cluster, and two nodes over real TCP sockets produce a trace matching the single-node run on both nodes — tensors compared by bit pattern, and the router's top-k compared exactly, because I3 makes that a claim of its own. It is a branch and not a second forward: block was split into sharedPart + combine so there is exactly one implementation of each, and the split was verified numerically neutral with the existing suite unchanged and green. The router runs on every node and its decision does not travel — the dense backbone is replicated, and a decision that shipped would be a second source of truth. The lockstep is real and worth naming: each node blocks on its peers inside every mixture layer, so the two must make progress concurrently (the test uses two threads), which is exactly why the synchronisation budget matters for M3. A silent peer makes the forward throw rather than emitting a trace from the experts this node happened to own, with the policy's deadline reaching the socket through the forward. Still not two processes, and not two machines — the real model at two nodes does not fit on this 8 GB host sources/DatacenterEngine/Qwen3_5Forward.swift, ShardedForward.swift, docs/m2-decisions.md
2026-09-16 The transport is a socket now, not a socket pair. D22: TCPListener binds, listens and accepts, TCPTransport.connect connects under the policy's deadline, and a contribution exchange over TCP is bit-identical to the single-node forward on the fixture's real weights — with a bring-up handshake over TCP as well. The socket pair could carry frames but could never fail the way a socket fails: nothing bound, no port was busy, no connection was refused, and all three now exist as tests. One detail that only a real socket teaches: a blocking connect blocks for the kernel's timeout, which is minutes, so the run would look hung rather than refused and the timeout policy would never apply — the socket goes non-blocking for the connect, poll waits for writability, and SO_ERROR answers, because POLLOUT also fires when a connection failed. Two wrong diagnostics fixed on the way: a failed connect reported as cannotBind (naming the wrong operation), and accept /connect silently shadowing the global functions of the same name, which Swift refused to compile rather than resolving quietly. DC-008 now carries only the half it is named for — the per-hop measurement, which needs the testbed sources/DatacenterEngine/TCPTransport.swift, docs/m2-decisions.md
2026-09-16 Bring-up: a cluster refuses to start when its nodes disagree (D21), and says which agreement failed. Six fields are compared before a token — protocol schema, family, revision, geometry and the plan digest — and a refusal names the field, so it reads "node 1 disagrees about planDigest" rather than "nodes disagree". What is deliberately not compared is an install digest: in a sharded run each node holds the experts it owns, so the installs differ by construction and comparing them would refuse every correct run — the kind of check that is obvious to add and wrong to add. The declaration frame carries its own magic, because a declaration and a contribution frame share one connection and a codec inferring the type from the payload would eventually read a declaration as terms; two tests hold that boundary from both sides. The peer set is checked too: a duplicate, a stranger and a silent node are three distinct failures — the first version collapsed them and could index an empty array, caught by reading it back rather than by a test. Twelve tests over a socket pair, including a matching handshake verified in both directions. DC-042 closed; DC-008 now carries the socket a real machine binds sources/DatacenterEngine/ClusterBringUp.swift, docs/m2-decisions.md
2026-09-16 The shard plan is data, and the format makes the dangerous mistakes unrepresentable. D20: owners is a flat array indexed by expert id, so an expert cannot be owned twice (one slot each) and cannot be quietly unowned (the length is checked against the model's expert count). The obvious alternative — a list of experts per node — makes both representable and leaves a validator to catch them afterwards. Compatibility is checked at load: schema, family and expert count, so a plan for another model is refused before a run rather than surfacing as a missing term; a test writes a truncated plan because that is what an editor produces. Identity is canonical JSON plus a digest, so two nodes compare before starting. The distribution is now contiguous blocks rather than round-robin so a node's reads walk forward through the stacked tensor — and the bit-identity tests were not adjusted and still pass, which is D17's claim demonstrated by a refactor rather than by argument. A run driven by a loaded plan file is bit-identical to the single-node forward at 2 and 3 nodes. Also removed: an emptyNode check no test could reach (the shape check catches it first, and fewer experts than nodes is legitimate) — deleted rather than kept as a case that looks like coverage. DC-040, DC-011 and DC-050 closed sources/DatacenterEngine/ShardPlan.swift, ExpertOwnership.swift, docs/m2-decisions.md
2026-09-16 DC-043 closed: the failure rules are implemented, and two decorative-parameter bugs died on the way. D19 decides what a single-user cluster run does when a peer goes quiet, stalls mid-frame, or dies: a clean timeout is retried (safe because D17's order is canonical, so the same terms resent give the same bits), a mid-frame stop is fatal and not retried (the stream is desynchronised), a closed peer fails, and a missing term fails the run rather than producing a smaller sum that looks like an answer. Duplicates merge only when they are bit-identical — compared on bit patterns, because Float says -0.0 == 0.0 and NaN != NaN. Then the fun part: the ExchangePolicy timeout never reached the transport, so a test configured 60 ms, waited 30 s an attempt, and passed in 90 seconds — behaviour asserted, duration not. Wiring it through found the same bug again one layer up, in the test's own decorator, which implemented send/receive but not applyTimeout and took the no-op default. Fixed both: 0.191 s, and the test now asserts deadlines == [60, 60, 60] so the deadline is known to arrive rather than assumed to. A parameter that is not wired up is a lie that only a clock catches sources/DatacenterEngine/ShardExchange.swift, ContributionTransport.swift, docs/m2-decisions.md
2026-09-16 The wire protocol exists, and a two-node exchange is bit-identical to the single-node forward. D18: little-endian, length-framed, floats as IEEE-754 bit patterns so -0.0, subnormals and NaN payloads survive — a test feeds the wire all of them and compares bits on the way back, because every other test would pass a codec that quietly cleaned them up. Every length prefix, dimension and term count is refused on the number before anything is allocated (two tests hand it a 2^31 length and a 2^30 width). No checksum, deliberately — D15's lesson: TCP checksums, and the completeness guard plus the trace digest catch the rest, so a third check is not worth a millisecond a layer. The test runs the fixture's real weights through a genuine exchange over a socketpair: each node computes only its own experts' terms, sends them, checks isComplete, reduces, and both nodes land on the single-node bits. SO_NOSIGPIPE on every descriptor, because a dead peer killing the process presents as "the node vanished" rather than "the write failed". Open: timeouts, retries and failure semantics (DC-043), and the transport measurement (DC-008). Also caught: the codec refused my own test data — a 12-wide frame declared with a 1-wide term — which is the check working sources/DatacenterEngine/ContributionWire.swift, ContributionTransport.swift, docs/m2-decisions.md
2026-09-16 M2's central property is demonstrated on real weights: the sharded forward is bit-identical to the single-node one. MixtureOfExperts.experts is now a thin wrapper over expertContributions + OrderedReduction.accumulate, so one node and N nodes run the same code rather than two implementations that agree today. ShardedMixtureTests takes the fixture's real quantized weights, the real router's selection and 2 and 4 simulated nodes, and compares bit pattern by bit pattern; the real-model 35 B digest is unchanged (b0d382db…), so the refactor is numerically neutral. The run forced two design additions, one because the first version was wrong: ExpertWeightProvider.serves(_:) so a node can decline an expert, and a production completeness guard (OrderedReduction.isComplete) because skipping is only safe if something checks the reduction saw every selected term — the first attempt failed loudly with "node 1 was asked for expert 6, which it does not own". The cost is measured and accepted: ~1.4 s added to a 16.3 s forward, paid rather than keeping a second fast path that could accumulate differently. What remains is the transport — nothing yet has crossed a network (DC-008, DC-009) sources/DatacenterEngine/OrderedReduction.swift, ExpertOwnership.swift, ExpertProvider.swift, docs/m2-decisions.md
2026-09-16 M2 starts, and its central contract is decided before the code: reduce contributions, never per-node partials (D17). M2's gate is "2 nodes bit-identical to the 1-node baseline", and fp32 addition is not associative, so the obvious design — each node sums the experts it owns and the totals are added — cannot reach it. Measured, not asserted: in fp32 at magnitude 2e7, (a+b)+c = 20000008 while a+(b+c) = 20000006 for a=2e7, b=c=3. So the all-reduce carries per-(token, expert) contributions and sums them in (token, expert) ascending order, which is exactly the sequence the single-node engine already performs. This is stronger than the brief's "fixed ring order": renumbering the ring, replacing a node or re-partitioning the experts cannot move a bit, so what must be pinned is the ownership map (a data artifact, DC-011) rather than the network's shape — which resolves the caveat left open in the brief response's I2. The cost is stated rather than hidden: ~64 KB per token per layer instead of an 8 KB partial (only the other node's terms cross the wire), still latency-dominated, with an exact-accumulator escape hatch if DC-051 ever says otherwise. OrderedReduction + four tests land with it, including the trap itself. DC-005 closed; DC-040/DC-041 now carry the planner and the transport docs/m2-decisions.md, sources/DatacenterEngine/OrderedReduction.swift, tests/DatacenterEngineTests/OrderedReductionTests.swift
2026-09-16 M1's gate passes, re-established with current evidence. The correctness claim was re-verified after D15/D16: the contract was re-run today and trace_diff reports IDENTICAL — 83 tensors, 0 elements, 40 discrete decisions, both digests b8c976c5e7ba8816… — the same digest as the original gate run, which is the evidence that the two performance fixes changed no arithmetic and that D11's flush is install-only. Generation is 0.108 tok/s cached (9.25 s/token) and 0.0374 uncached, 5.6x the historical 0.0191, and both paths generate the identical tokens. Peak memory is 348.6 MB — 12x below the 4.16 GB the checkpoint run reported, because the install path reads uncached and maps nothing. Three claims in the gate doc were corrected rather than edited away, including "M1 has no KV cache": --cached exists and three tests hold it to the sequence path's numbers. A resource finding too: the checkpoint contract step drove swap to 402 MB free on this node before being stopped, so it is a heavy job that must run alone. New open item DC-108: no Python contract reads an install docs/m1-gate.md, DC-108, README.md, AGENTS.md
2026-09-16 The instrument lied again, and fixing it gave another 24%: 39.9 s → 15.0 s, 2.67× overall. Reading the whole-tensor path for DC-107 found that digestMatches verified regardless of the verify flag — reading five rows of the 1.02 GB embedding read all of it to hash it, which is the embed phase at 1.64 s for kilobytes of data — and that those reads were never counted, so the "measured 3.06 GB" was use-traffic reported as a total. Third instrument artefact in this project, after an elementsRead * 2 estimate and a cache-assisted benchmark. payload also read every entry twice, once to hash and once to return. D16: verification and use share one buffer, verifyOnFirstUse is explicit and off by default (integrity stays out of band), and SourceTiming.verifiedBytes exposes the verification share. Measured on the real model: 19.8 → 15.0 s, embed 1.64 → 0.00 s, head 3.20 → 1.64 s, load 2.95 → 1.74 s, digest unchanged (b0d382db…). Prefill is now 0.321 tok/s against 0.127 two rounds ago, and 3.06 GB is finally the whole truth rather than the part that happened to be counted (DC-107 for what remains: attn.core 3.86 s, unpack 3.45 s, reads 2.99 s) sources/DatacenterEngine/Install.swift, docs/m1-decisions.md, docs/m1-gate.md
2026-09-16 Profiled the forward and found a 21 s bug: a 2x speedup, same digest. DC-105 asked where a five-token forward's 39.5 s went, so SHARD_PROFILE=1 now times the real path and the phases sum to the run (39.82 s of a 39.9 s wall clock). mix.read — the expert fetch — was 64%, and splitting it open inside the reader showed SHA-256 slab verification on every read at 21.04 s, 53% of the whole forward, against 0.82 s of reads and 4.06 s of unpacking. D15: verification is now a parameter, off by default, with integrity established out of band by quantize.py verify — the same guarantee without paying it on the hot path. Verified on the real model: 39.9 s → 19.9 s, mix.read 25.53 → 6.28 s, digest seconds 0, trace digest unchanged (b0d382db…). The hypothesis I opened last round was wrong and the profile says so: the dense weights re-read per forward are 15%, not the missing 3.5x (DC-106 stays open, smaller). Two independent confirmations fell out: the reads now measure 1.0 GB/s (3.04 s for 3.06 GB), reproducing the standalone cold measurement from the other direction, and the 3.06 GB is the measured install traffic of one forward — 2.02 GB of experts plus ~1.04 GB of dense. Writing the test that pins "verification off costs nothing" found a flaw in the instrument itself: the clock had started outside the branch and booked 4.2e-08 s of the branch test as verification. Prefill is now 0.251 tok/s against 0.127 (DC-107) sources/DatacenterEngine/ForwardPass.swift, Qwen3_5Forward.swift, MixtureOfExperts.swift, Install.swift, docs/m1-decisions.md, docs/m1-gate.md
2026-09-16 The read path is fine, and two of my claims are withdrawn because of it. tools/measure_expert_reads.py performs the engine's own fetch — six preads per expert, geometry from the install rather than assumed — over 3.10 GB, and reports 997 MB/s ascending, 979 MB/s with the expert order shuffled, against the SSD's 1161 MB/s sequential ceiling. Cold and repeated runs agree (997 against 989), so there is no cache effect: the reads are real and they are near the ceiling. The first version of the tool read 77.6 MB and reported 2957 MB/s — faster than the hardware — because it was measuring cache, the same trap the Testbed page already recorded. Withdrawn as a result: the "effective 177 MB/s" (computed from the elementsRead * 2 estimate) and the conclusion that the single node was running at a fraction of its SSD's capability with local read work as the first priority. Both were wrong, and the second was advice. What the numbers now bound: expert reads 2–4 s, dense re-reads ~3 s, dequantisation ~2.9 s, matmuls ~1.5 s — 9–11 s of a 39.5 s wall clock, leaving ~3.5x unaccounted for and not the disk's fault (DC-105, and DC-106 for the per-token dense re-read) tools/measure_expert_reads.py, docs/m1-gate.md, DC-105, DC-106
2026-09-16 The "bytes from SSD" figure was an estimate, and it overstated the traffic by ~3.5×; the measurement replaced it. datacenter-trace computed expertElementsRead * 2, which assumes bf16 on disk — but this install stores experts at 0.578 bytes per weight (4-bit codes with group-64 scales and zeros), so the number I published yesterday evening was wrong in the direction that flatters the project. The install had been counting the bytes it actually hands out all along, so WeightSource now exposes that counter, datacenter-trace reports install_bytes_read_this_forward, and three tests hold it — including the one that matters beyond the fix: a forward reads a slice of the install, not the file, which is the streaming claim and the same class of bug as DC-088. The sweep's real traffic is bounded by geometry to 404–807 MB per prefill token (one expert is 1,818,624 bytes; the fetch count is Q9), and the derived "effective 177 MB/s" reading is withdrawn — it was computed from the estimate, and a number about the read path has to come from the read path sources/DatacenterEngine/Install.swift, ForwardPass.swift, DatacenterTrace/main.swift, tests/DatacenterEngineTests/SourceBytesTests.swift, docs/m1-gate.md
2026-09-16 The cache is transparent, its hit rate is structurally zero, and a prefill token costs 1.4 GB of SSD — all three measured on the real model. The authorised sweep ran capital's five tokens through the engine at bank sizes 2, 8 and 16: three traces, byte-identical, checked independently rather than inferred — trace_diff reports 83 tensors, 0 differing elements, 40 discrete decisions and matching digests (b0d382db…). So a bank's size cannot change what the engine outputs, which is the least a cache has to be, and the same property M2 will need from sharding on the same harness. The expert counters are identical at every size: 6,977,224,704 bytes from SSD for five tokens (1,395 MB per token), 2,218 requests, 0 hits, 13.95 GB decoded into memory (2× the SSD bytes: int4 → fp32). Prefill took 39.3–39.9 s for the five tokens (7.90 s per token, 0.127 prefill tok/s) — a prefill rate, not the gate's generation baseline, and the two must not be conflated. Two things this settles and one it exposes: the hit rate is 0.0000 regardless of size because the banks are rebuilt per forward, so D12 is not "how big" but "may a bank survive a token, and at what total budget"; the run pushed swap from 935 MB to 1,504 MB, so the sweep's own 1 GB guard now refuses a second run until the machine settles; and the digest moved off the pre-D11 b8c976c5…, exactly as D11 predicted, so the contract comparison under the flush is still M1's outstanding half. One figure is recorded unexplained rather than guessed: 2,218 requests against 320 selections per token .build/m1-sweep/, docs/m1-gate.md, DC-034
2026-09-16 The tracker became a tracker: open items only, and the record moved here. It had grown to 82 task rows of which 43 were Done, plus a 200-line dated log inside a page whose stated job is what is not finished. The closed rows are gone, each phase keeps a one-line index of what closed, and the dated record lives on this page — because a live list that is 43 rows of finished work deep is one nobody reads. Corrected in the same pass: the tracker's own header still claimed no source code exists at all, its status row still quoted the retired 0.0191 tok/s, and its counts disagreed with its own rows .wiki/Project-Tracker.md, .wiki/News.md, DC-104
2026-09-16 DC-087's divergence is resolved by D11, and re-checking it was worth a round. The task was filed as "Metal flushes denormal operands, so no GPU kernel can be bit-identical for free". D11 then made the CPU flush the same operand — the same threshold, by the definition of what denormal means — so the two should now agree on the operand itself rather than on a tolerance. Re-checked by restoring the full grid and removing both skips: on the shapes that exposed the divergence in the first place, the partly-filled groups (columns = 5 and 65, group = 4 and 64), the GPU and CPU now agree on every value — 0 differences out of 130. So the specific fault this task was opened for is gone, and the honest restatement of D10 is narrower than it read: a GPU kernel is not ruled out by denormals once the contract defines the flush, and the remaining question about Metal is the 1.30× against a CPU path that is already bit-identical. What is not resolved: the full grid still fails on some shapes — group = 1 cases and columns = 128, group = 4, rows = 3 — which is a different cause from the denormal one and is not diagnosed, so it is recorded as the next probe rather than guessed at. The grid keeps its skip until then, and the tree is green: 97 Swift tests, 139 Python tests/DatacenterEngineTests/MetalUnpackTests.swift (reverted to its skipped state), DC-087, docs/m1-decisions.md (D10, D11). 97 Swift tests, 139 Python
2026-09-16 DC-088's mystery is solved, and it was two checks rather than one. Reading the region properly, instead of patching it from memory, shows InstallFile.rows carried two whole-entry digestMatches checks on the int4 path: one at line 205, before the three payload reads, and one at 225, after them. Every previous attempt replaced or removed exactly one — my patches matched the second, my deletions targeted the first — so a whole-entry check always survived, and it rejected any tamper anywhere in the entry before a slab was ever consulted. That is precisely why a test that tampered the payload's last byte kept reporting the first expert as bad, and why I twice concluded the offset arithmetic was wrong when it was not. With both checks removed and the corrected guard in place the suite still fails, and this time with a crash rather than an assertion, so there is a fourth fault. My first guess at it was wrong: the crash surfaced while a GatedDeltaNetTests case was running, not the new cold-read test, which means the fault is on an Install path the suite exercises in some other test and the new test is not the obvious suspect. Six rounds, three faults identified, one fixed and measured, the tree green after each. The remaining work is now a short list rather than a mystery, and it wants a session that starts by instrumenting the test as well as the engine DC-088, tools/quantize.py (writer, tested and proven), sources/DatacenterEngine/Install.swift (reader, reverted). 95 Swift tests green, exit 0
2026-09-16 DC-088 parked, with three located faults and an honest reason. The instrumented run — the thing my own note said to do three rounds ago — crashed with signal 5 instead of failing an assertion, which means the slab slicing reads outside its buffer for some request; before the crash it printed nothing, so the crash happens on the first read. Together with the two faults already found (the guard testing range.count % inner == 0, which is 1 % 32 for one expert of a stack and skipped the branch entirely; and a whole-entry digestMatches check sitting in front of the slab branch on the int4 row path, one of five call sites in the file), that is three independent defects in about forty lines. Five rounds have gone to this task, each ending by restoring the committed state because a red tree is worse than an unfinished change, and each making real but partial progress. The judgement I am recording: this does not need another fragment of a round, it needs one session that instruments the read path first, characterises each defect with a print, and only then edits. Until there is room for that, the cost it removes — tens of gigabytes of extra I/O per process on the real model — stays on the books as an estimate DC-088, tools/quantize.py (writer, tested and proven), sources/DatacenterEngine/Install.swift (reader, reverted). 95 Swift tests green, exit 0
2026-09-16 DC-087 narrowed from a guess to a shape. The previous round guessed the GPU/CPU divergence was a partly filled final group; a diagnostic grid of twenty-four shapes says that guess was wrong. Exactly one shape fails — columns = 64, group = 1, rows = 2 — with the first differing index at 117 (row 1, column 53, group 117), where the GPU produced bits 0 and the CPU a small non-zero value. Single-row versions of the same shape are identical, and so are two-row shapes at groups 4, 8 and 64, which rules out both "many groups" and "many rows" as triggers on their own. Twenty-three of twenty-four shapes agree, so the kernel is right about the common path and wrong about one corner — which is the more useful statement than "the GPU does not work", and it is why the diagnostic prints rather than asserts. Next: dump the shader's own view of the group index and the scale for that index, since reading both implementations has now failed twice to explain it tests/DatacenterEngineTests/MetalUnpackTests.swift (testWhichShapesDisagreeDiagnostic), DC-087
2026-09-16 ExistentialAny enabled, and the reason it looked expensive was a measurement artefact. The register said fourteen warnings and deferred it. The real count is four any keywords in one file — and the mistake was not the arithmetic: an incremental build does not re-emit warnings, so the fourteen came from an output that happened to carry unrelated diagnostics, and every count taken afterwards came from a warm build that printed nothing, which read exactly like a clean tree. Touching the sources first produced the true 7 lines / 4 unique sites, the fix-its were applied at the compiler's own column positions, and swift test with the feature on now exits 0 with zero diagnostics and 95 tests green. shardLanguageStandard therefore carries five upcoming features on all six targets. This is the second cost in that register to be wrong in the same direction — first overstated (MemberImportVisibility at "expensive", actually one import), now understated — and both were guesses wearing numbers, so the method section now begins by touching the sources Package.swift, docs/swift-language-standard.md, sources/DatacenterEngine/MetalUnpack.swift, tracker DC-015. 95 Swift tests green
2026-09-16 Eight codes per load, and the unpack is now 1.83× the scalar. The nibble extraction the previous round left unmeasured: unpacking from a UInt32 (eight codes) instead of a UInt16 (four) gives 1185.5 M values/s against the scalar 648.3, where four-wide gave 1038.3 — so the wider load is worth a further 14 %, and a 35 B token's unpack falls from 3.3 s to 2.9 s. The arithmetic per element is unchanged (one multiply, Float(code - zero) * scale), so bit-identity still holds by construction, and the sixty-shape grid plus the end-to-end golden tests confirm it. The guard matters: a wider block must not straddle two groups, because the scale and the zero point change at the boundary, so the wide path runs only when the group is a multiple of eight and the four-wide loop stays as the general case. What is left here is thin — the widest useful floating-point lane is 128 bits and Swift offers no cheap widening from integer lanes — so the next real factor is Metal, not more of this sources/DatacenterEngine/Install.swift, tests/DatacenterEngineTests/Int4UnpackTests.swift, docs/m1-decisions.md. 90 Swift tests green under -Onone and -O
2026-09-16 D11: the denormal rate is 0.72 %, measured, and it decides the Metal question. tools/measure_denormals.py reads the scale sections of an install uncached and counts: on the real 35 B install, 3,781,952 of 523,304,960 scales are denormal (0.722705 %), none zero, none non-finite, read in 2.13 s with free disk steady at 17 GB. They are not spread evenly — layer 0's gate_up_proj is 41 % denormal by itself, layer 10 has 64 — which is what near-zero weight groups in the first layers look like. Four million differing weights per pass is not a rounding detail when the gate compares every value, so there are two options and no third: flush denormals in the contract (round denormal scales to zero in the quantizer and the engine, so both sides agree because the flush is defined; the lost values are ~1e-38 and I3 is untouched because routers are bf16), or keep Metal off the entire expert path, since a flushed scale makes an unpacked weight zero where the CPU makes it denormal and that propagates into every matmul. I recommend the first — it is a definition rather than an error, one line on each side, and testable the same way as everything else — but it changes what bit-identical means, so it is recorded for a human decision rather than taken unilaterally tools/measure_denormals.py, docs/m1-decisions.md (D11), DC-087. Measurement only; no engine code changed
2026-09-16 The first kernel change: reading one expert no longer decodes the stack. InstallFile.rows(named:range:) on an int4 tensor used to call tensor(named:) and slice the result — for the real model's [256, 1024, 2048] expert tensor that is 537 M parameters and two gigabytes of Float on a node with four and a half, to return one expert. The payload is laid out section-major (every row's codes, then every row's scales, then every row's zeros), so each section is row-contiguous and a row range is three bounded reads and one decode. Measured on the fixture: one expert of an eight-expert stack now reads 1184 of 9472 bytes, exactly its one-eighth share, against the whole 9472 before. Five tests pin it: every expert of every stacked fixture tensor equals its slice of the whole decode, rank-2 row reads are unchanged, the byte cost is measured rather than asserted, the whole-tensor read still returns every value, and out-of-range is refused. And the bug the tests caught is the lesson: the first version used the range as a payload-row index, so range 1..<2 meant "the first row of expert 1" rather than "expert 1" — the fixture's gate_up_proj returned 32 values where 1024 were expected, and every shape was plausible. The range is in leading-axis entries, always, and has to be translated by the rows-per-entry factor before any payload arithmetic. That trap has now bitten four times in this project — both contract readers, the install builder's chunking, and here — so it is written down as a rule rather than remembered sources/DatacenterEngine/Install.swift (dequantizeInt4 takes a row count; the int4 row read does three bounded reads; bytesRead makes the cost observable), tests/DatacenterEngineTests/InstallRowReadTests.swift. 82 Swift tests, 136 Python
2026-09-16 D11 option 1 endorsed, and the implementation immediately taught me where the switch does not go. The operator chose the flush — "one line on each side versus the whole hot path staying on the CPU" — so the direction is settled: denormal scales are read as zero, in the Swift unpackers and in the Python reader, so the two agree because the flush is defined rather than discovered. Two traps surfaced on the first attempt, both worth recording because they are the whole difficulty of a one-line change. The quantizer must not flush the scale before deriving the zero point: the zero point is QMIN - minimum / scale, so flushing first divides by zero and the Python suite failed with it. The flush belongs after the zero point, or in the readers only, which is where agreement actually has to hold. And the two Swift unpackers hold the scale in different types — the vector path as a Float, the scalar path as the UInt32 word it was loaded as — so one textual patch matched the wrong one and the build failed on SIMD4<Float>(repeating:). Both attempts were reverted rather than left half-applied, so the tree is green: 96 Swift tests, 137 Python. The change is small, the plan is now exact, and it has its own task DC-089, docs/m1-decisions.md (D11). No code changed in the end
2026-09-16 DC-015 finished properly: the fourth feature cost one import, and both my estimates of it had been guesses. MemberImportVisibility was first declared free on a swift build measurement that never compiled the test targets, then declared expensive on a pass I interrupted. Neither was a number. Iterating swift test under the flag and adding exactly the imports the compiler named converged in one round with one import — Qwen3_5ForwardTests.swift needed import DatacenterIR — and swift test now passes with the feature enabled, so shardLanguageStandard carries four upcoming features on all six targets and the tree is clean at zero diagnostics. The register has a cost table instead of adjectives, the method section says swift test and says why, and it carries one more trap that cost me a crashed loop: this repository's path contains a space, so a diagnostic-path regex that stops at whitespace truncates it and the loop then edits a file that does not exist Package.swift, docs/swift-language-standard.md, tests/DatacenterEngineTests/Qwen3_5ForwardTests.swift, tracker DC-015. 95 Swift tests, exit 0
2026-09-16 The first Metal kernel exists, is measured, and is deliberately not used yet. MetalUnpack compiles the shader at runtime with mathMode = .relaxed (fastMathEnabled is deprecated), dispatches one thread per output over the three payload sections, and is checked against the scalar op bit for bit — and that check found a real divergence: with columns = 65 and group = 4, so that a row's final group holds one real value of four, the GPU differs from the CPU on exactly that last column of every row and nowhere else. Reading both implementations shows the same expression with the same inputs, so the next step is to dump the shader's view of the group index and scale for that column rather than to guess. Until it is fixed the kernel is not called by anything, which is what D10 says a kernel must earn — and the test is skipped with that reason rather than deleted, filed as DC-087. Where it does agree it is modest: on a well-formed 8192 × 2048 payload with group = 64, 1178 M values/s on the CPU against 1567 M on the GPU, 1.33× — so the GPU is not obviously the answer for this op either, and that measurement is the reason not to wire it in on optimism sources/DatacenterEngine/MetalUnpack.swift, tests/DatacenterEngineTests/MetalUnpackTests.swift, DC-087
2026-09-16 DC-087's residue: two diagnostics, two failures in my own tooling. The residual divergence — group = 1 shapes and columns = 128, group = 4, rows = 3 — is undiagnosed, and the reason is worth recording rather than hiding: both attempts to instrument it failed before reaching the thing being diagnosed. The first print block used a Swift multi-line string inside a shell heredoc and printed its own format text; the second, written by hand with string concatenation to avoid exactly that, did not compile (String(Array(...)) does not mean what it looks like, and a new test file cannot see the private helpers in MetalUnpackTests). Both were reverted, the tree is green at 97 Swift tests and 139 Python, and the task carries a sharper instruction than it had: make the diagnostic a minimal edit to a test that already compiles, which is precisely how the slab instrument finally worked after four failures of the other kind. The D11 half of this task — the part that was filed — is genuinely resolved and proven tests/DatacenterEngineTests/MetalUnpackTests.swift (reverted), DC-087. 97 Swift tests, 139 Python
2026-09-16 DC-087 is fully diagnosed, and the answer is a sign bit. After D11 resolved the denormal half, the residue resisted three rounds — and the diagnosis, once the instrument finally ran as a minimal edit to a compiling test, is that every remaining difference is +0.0 against −0.0: 0x00000000 on one side, 0x80000000 on the other, five cases across a forty-shape grid, with cpuZero and gpuZero both true and no NaNs left. Part of the original residue was a different artefact of the same hunt: the synthetic payloads are random bytes, so about 0.4 % of their scale values are NaN, and two NaNs with different payloads are unequal by bit pattern while being equally NaN — the real install has 0 non-finite scales, measured, so that half was the fixture's fault rather than the kernel's, and it accounted for one of the six failures. What is left is not a bug to fix but a question to answer, filed as DC-090: does bit-identity distinguish the sign of zero? The GPU normalises some zero results and the CPU does not, so either the contract defines a rule — consistent with D11, which defined the denormal flush rather than living with the difference — or the GPU path cannot be bit-exact on this grid. It is the operator's call, not a test convenience, so the comparison was reverted rather than quietly made zero-aware. Tree green at 97 Swift tests and 139 Python tests/DatacenterEngineTests/MetalUnpackTests.swift (reverted), DC-087, DC-090. 97 Swift tests, 139 Python
2026-09-16 Auditing my own fix one round later found a 8.05 GB mistake, and the audit step earned its place. DC-091's fix — persist the expert slot banks so tokens can hit them — was made, tested at fixture scale, and shipped with 98 green tests. The audit then asked the question the fixture cannot: how much memory is that? One expert is 12.58 MB of fp32 (8.39 gate+up, 4.19 down), expertSlotsPerLayer = 16 makes that 201 MB per layer, and keeping all 40 layers resident is 8.05 GB against roughly 4.5 GB usable. On this node that is not a slow path, it is the swap configuration that panicked the machine twice. The fix was reverted, which is the correct use of a revert: the change was not wrong about the hit rate, it was wrong about the budget, and the brief's own "per-layer LRU slot banks" never said how large. 97 Swift tests, back to the pre-change count. What this leaves is a decision with numbers attached, filed as DC-092/D12: a 1.5 GB cache across 40 layers is 2.98 slots per layer, so the code's 16 is about ten times too large, and the split between slots and layers is a real trade — slots buy cross-token hits, layers buy nothing at all once a bank is dropped, and with top-k 8 against 256 experts a one-slot bank will still hit rarely. That question is the operator's, because it trades RAM against hit rate on a node whose RAM limit has already caused two panics docs/m1-decisions.md (D12), DC-091, DC-092, revert 7a902ae. 97 Swift tests, 144 Python
2026-09-16 The toolchain became a gate, and writing the gate found the bug that made the old check a no-op. The operator's rule is Xcode 27 with Swift 6.4 and no exceptions, which the Swift CI job did not honour: it printed ::warning:: and skipped when the runner was below the floor, so a green run could mean the toolchain was right or the toolchain was so wrong that nothing ran. tools/check_toolchain.py now parses both version outputs and refuses anything that is not the standard — older and newer, because a newer Swift is a project decision rather than drift — and the CI job has no skip branch: it fails until a runner image carries Xcode 27. That job is therefore expected to be red today, and a red Swift job means the runner image is below the standard, never that the code is broken; that sentence is what the job itself prints. The bug worth the whole change: the step being replaced read version=$(swift --version 2>&1 \| tail -1) and matched Swift version 6.4 against it. The last line is Target: arm64-apple-macosx27.0.0, so the pattern could never match and the skip branch ran on every runner the project has ever had — including one with the correct toolchain. A gate that reports success by not running, hiding inside a warning. Eleven tests now pin the parsing, both refusals, the constants, and that tail -1 specifically .github/workflows/swift.yml, tools/check_toolchain.py, tools/test_check_toolchain.py, README.md, AGENTS.md (conventions, traps, build-and-run), CONTRIBUTING.md, .wiki/Home.md, DC-103. 105 Swift tests, 174 Python
2026-09-16 DC-088 has two faults, not one, and the second is why the first kept looking wrong. The slab branch guarded on range.count % inner == 0 — 1 % 32 for one expert of a stack — so it was skipped and every read fell back to the whole-entry digest. Fixing that guard was measured working. But Install.swift also carried five digestMatches(entry) call sites, one of them on the int4 row path immediately before the slab branch, so any tamper anywhere in an entry failed the whole-entry hash before a slab was ever consulted — which is exactly why a test that tampered the payload's last byte still reported the first expert as bad, and why I twice concluded the offset arithmetic was wrong when it was not. Removing the duplicate check alone did not turn the suite green either, so at least one more thing is involved and I did not find it. Fourth round on this task; each one ended by restoring the committed state, which is green at 95 tests, rather than leaving a red tree for a nearly-finished change. The method note for whoever picks it up: stop reading this code and instrument it — print, for one read, the branch taken and the digest computed, which is the technique that eventually found the Metal divergence after three failed readings DC-088, tools/quantize.py (writer, tested), sources/DatacenterEngine/Install.swift (reader, reverted). 95 Swift tests green, exit 0
2026-09-16 The wiki's tables were gated only when I remembered to run the tool. .wiki/ is a separate repository and is gitignored, so check_markdown_tables.py's default globs find it on this machine and find nothing on a runner — which means the document the project calls its working record, and where the tracker lives, had no gate at all in CI. Three misaligned wiki rows this session were caught by a command I typed by hand, which is the same failure the tool was built to end, one level up. The Markdown links job now clones the public wiki and checks its tables, with the fallback this project's release rules require: if the clone fails, the wiki is reported NOT CHECKED with a ::warning:: rather than passing quietly. The link gate is deliberately not run there, because wiki pages link to each other by page name ([Roadmap](Roadmap)) and the checker would call every one of them broken — a gate that cries wolf is worse than no gate. Verified locally first: the tool reports 9 wiki files, 0 misaligned when given those paths explicitly, which is exactly how CI passes them .github/workflows/markdown-links.yml, tools/check_markdown_tables.py, DC-095. 97 Swift tests, 150 Python
2026-09-16 I6's missing half is fixed and verified against an independent hash. The audit found source.files empty in the real install and both committed fixtures, so a converted model could not name the weights it came from. quantize.digest_snapshot now digests every source *.safetensors through the uncached descriptor — one window at a time, because the node has already lost 14 GB of free disk once to a page-cached read — and the build records the map, so the next install of any model carries it. The check is against a separately computed sha256, not against the code path having executed: {"model.safetensors": "e9a9150e…"} matches hashlib.sha256 of the file, and a digest that is present but wrong would have satisfied the weaker test. Cost is named rather than hidden: this is a second pass over the shards, tens of seconds at the measured 1161 MB/s on the real model, which is worth paying because the alternative is an artifact that cannot say what it is. Still open, and now precisely: revision needs the commit that a build does not know, and the committed fixtures must be regenerated so they stop carrying {} tools/quantize.py, tools/test_quantize.py, DC-098. 100 Swift tests, 153 Python
2026-09-16 I filed a false finding, and it reached the public gate documentation before I caught it. DC-093 claimed the MoE fixture never exercises partial RoPE. It always has: the fixture's config.json carries "partial_rotary_factor": 0.5 and "rope_theta": 10000000.0 at the top level, and spec.json carries "partialRotaryFactor": 0.5. My check read the config under text_config — because that is where the real checkpoint nests its geometry, which I had fetched the round before — so it printed (absent) for keys that were present, and I reported my script's blind spot as a fact about the repository, complete with a table and a new tracker row, and added a section to docs/m1-gate.md claiming an untested path. Both are now withdrawn and the gate doc carries the withdrawal where a reader will see it. This is the fourth time this session that an instrument of mine lied and I believed it: the heredoc that printed its own format text, the print block that would not compile, the group = 1 probe that measured my own arithmetic, and now a JSON lookup in the wrong nesting. The pattern is not that the repository is wrong; it is that I keep reporting tooling output as evidence without checking the tool. The counter-measure is the one that has actually worked all session — a minimal change to something already proven to work — and it is now written down as the rule I follow when a finding is about to be filed docs/m1-gate.md (withdrawn claim), DC-093 (withdrawn). 97 Swift tests, 144 Python
2026-09-16 The sweep instrument's mechanism is pinned, and the first hit-rate curve exists. The gate wants a measured cache hit rate and D12 wants a swept one, so before asking for a real-model run I checked the thing the sweep depends on: an instrument that has never been run is not an instrument. ExpertSlotCacheTests drives a counting stub and establishes three promises — the bank is bounded (peakResidentExperts <= 2 for a bank of two, which is the figure the byte budget multiplies), eviction is LRU as the brief's own words require (touch 0, 1, 0, then 2, and 0 survives while 1 dies), and the hit rate is non-decreasing in capacity, or a curve could not be read. The curve for one skewed 24-step sequence over eight experts: 1→2/48, 2→18/48, 4→26/48, 8→32/48. My first assertion said the bank of eight must hit all 48, and the real answer is 32 — eight experts times two projections is sixteen compulsory first misses. The engine was right and my arithmetic was wrong, which is this session's recurring shape and the reason the test is worth having: a fixture this small says nothing about the real model's routing, but it says the mechanism honours its contract, so a real curve will mean something tests/DatacenterEngineTests/ExpertSlotCacheTests.swift, docs/m1-gate.md, DC-092. 105 Swift tests, 156 Python
2026-09-16 The failure that cost a round became a tool, and the tool immediately caught the same mistake being made again. DC-093's false claim came from a hand-written lookup that knew one config nesting; the answer is not to be more careful but to make the comparison a check with tests, like the link gate and the table gate. tools/test_fixture_spec_matches_config.py asserts that every geometry field in a checkpoint's config reaches the spec with the same value and that nothing is silently defaulted — because a default looks exactly like agreement, and this project already treats a missing policy role as a stop rather than a default. The first version of the tool looked in the wrong place again: the fixture keeps rope_theta and partial_rotary_factor under rope_parameters, which is also where the real checkpoint keeps them and what the importer reads (partialRotaryFactor: rope?.partial_rotary_factor), so it reported both as absent from a config that has both. This time the tool printed the mismatch, I checked the source, and the false claim never reached the tracker — one round instead of a retraction. It now reads three nestings, has six tests including the one its author got wrong, and with it the answer is clean: the real install's spec.config carries all 23 geometry fields of Qwen/Qwen3.6-35B-A3B at the right values, partialRotaryFactor 0.25 among them, so the validation model runs the arithmetic the checkpoint declares tools/test_fixture_spec_matches_config.py, docs/m1-gate.md, DC-094. 97 Swift tests, 150 Python
2026-09-16 D10: the kernel contract for Metal, decided before the first kernel is written. D9's rule generalises to the GPU as one accumulator per output, k ascending — which does not exclude Metal: a one-thread-per-output kernel satisfies it trivially and a shared-memory tiled kernel satisfies it too, since a tile contributes a contiguous run of k values. What is excluded is split-K and reassociation, which is exactly what a fast GEMM does because it is fast. The trap is measured rather than assumed: Metal compiles shaders with fast math enabled by default, and fast math is the licence to reassociate and to fuse a * b + c into an FMA — so a kernel with identical source arithmetic can produce different bits unless it is compiled with fastMathEnabled = false and checked against the scalar op bit for bit. Probed on this node: MTLCreateSystemDefaultDevice() returns an Apple M2 with unified memory, and a runtime-compiled MSL shader compiles. And a gate rule falls out of it: the GitHub runner has no GPU, so every Metal test must skip when there is no device — the same shape as the Swift-6.4 skip in DC-036, which has been bitten twice, applied before it costs anything docs/m1-decisions.md (D10). No engine code changed this round; the probe was a throwaway program
2026-09-16 The table gate is now a tool, because I made the same mistake twice. Two tracker rows in a row landed in the three-column story table below the task table — a five-column row in a three-column table renders as garbage — and both times the flaw was found by a shell one-liner run by hand, not by the repository. tools/check_markdown_tables.py now does what that one-liner did: every run of consecutive table rows must agree on its column count, with escaped pipes and inline code spans excluded because Markdown excludes them. It ships with seven tests, most of them about the failure case, because a checker that cannot fail is a checker that lies — and the tests immediately earned their place by catching the tool's own misleading message: it counted pipes and called them columns, so a five-pipe row was reported as five columns when it has four. The tool was fixed rather than the tests. It runs in CI beside the link gate, so the next one is caught by the repository instead of by me tools/check_markdown_tables.py, tools/test_check_markdown_tables.py, .github/workflows/markdown-links.yml, tracker. 144 Python tests, all stdlib, all run by CI
2026-09-16 The run I keep asking for became a script that refuses to start. Everything the sweep needs existed but lived in my messages rather than in the repository, which is the wrong place for a heavy procedure on a node that has panicked twice. tools/run_m1_sweep.py runs the same prompt at several bank sizes, collects each run's metrics into a report and a printed table, and — the part that matters — checks the two conditions that preceded both panics before doing anything: the 5 GB disk floor through require_headroom, which also honours the watchdog's stop marker, and swap already in use, reported always and refused above 1 GB. A machine well into swap is not the machine to start a 35 B run on, and that is now code rather than a sentence. --dry-run prints the plan and executes nothing. The metrics file's schema is read, not assumed, and its extraction walks in document order — the first version used a LIFO stack, which reversed "the last layer's figure wins", and its own test caught the reversal with 3 != 4. That is the second time this session a test I wrote found my arithmetic rather than the engine's, which is the point of writing them before believing anything tools/run_m1_sweep.py, tools/test_run_m1_sweep.py, docs/m1-gate.md, DC-100. 105 Swift tests, 162 Python
2026-09-16 D11's flush is implemented, in four places, with the placement the whole difficulty. The operator chose the flush — denormal scales read as zero, on both sides, so CPU and GPU agree because the flush is defined — and DC-089 lands it: a flush_denormals helper and its use in the Python reader and on the quantizer's stored scale, and a flushed helper used by both Swift unpackers. The trap that cost the first attempt is worth writing down: quantize_rows divides by the scale to pack the codes and to derive the zero point, so flushing early divides by zero and the Python suite said so immediately; the flush belongs on the value that finally gets stored, after the codes are built. The other trap was a type: the Swift vector unpacker holds the scale as a Float and the scalar one as the UInt32 word it was loaded from, so a single textual patch matched the wrong one. Proven by tests the fixture could never provide — it contains no denormals at all — with a crafted denormal scale decoding to zero on both Swift paths and in Python, the two agreeing bit-for-bit, normal scales keeping both value and sign, and the quantizer storing zero. One consequence stated rather than discovered: the real install's stored scales still contain denormals, every reader now flushes them, so the engine and the contract still agree with each other while both differ from traces recorded before this change — capital's digest b8c976c5e7ba8816… is now historical tools/quantize.py, sources/DatacenterEngine/Install.swift, tests/DatacenterEngineTests/Int4UnpackTests.swift, tools/test_quantize.py, docs/m1-decisions.md. 97 Swift tests, 139 Python
2026-09-16 I finally delivered the thing the brief actually asked for. Re-reading the objective rather than my own summary of it: it ends with "do NOT start coding — respond with (a) your understanding of the invariants, (b) any place where this brief is underspecified or where you disagree, and (c) a concrete plan for Milestone 0 only." That deliverable had never been written in its measurement-grounded form, and it needs no approval, so it should have come many rounds ago. docs/brief-response.md now carries it. What the work changed about (a): I1 is not mainly about floating-point determinism — the arithmetic was right every time it was asked precisely, and the four things that actually threatened reproducibility were machinery (a bank rebuilt per forward, an undefined denormal flush, NaN payloads in my own fixture, a memory budget that would have swapped the node). What (b) now contains, each with a measurement: the slot bank's missing size (8.05 GB against 4.5 GB usable), the 16 KB/O_DIRECT coupling that makes the alignment requirement deferrable rather than unmet, the "4–10×" target being a read fraction rather than a range, and the all-reduce being serialisation-limited at the measured 0.49–0.64 ms RTT rather than bandwidth-limited — which makes overlap a requirement, not an optimisation. Plus one gap the brief leaves: the reference build needs pinning, not only the source revision. And (c) is M0's state, including the caveat the brief would not have known to ask for: M0's gate is proven at fixture scale, because this node can hold a fixture and not a 1B model docs/brief-response.md, DC-101. 105 Swift tests, 163 Python
2026-09-16 M1's cache hit rate had a structural cause, not a statistical one — and it is fixed. The gate measured 0 hits over 2240 requests, which reads like a model that never revisits an expert; the cause was simpler and worse. Qwen3_5Forward.loadLayer built a fresh ExpertSlotCache on every forward, and a forward runs once per token, so every token began with an empty bank — a bank dropped with the layer is not a bank, and the brief asks for per-layer LRU banks. The banks now live in a small reference type held by the forward, created per layer on first use and reused afterwards, while the rest of a layer's weights keep their load-and-release lifetime (the 3.2 GB stacked expert tensor is still never materialised). Proven at fixture scale, with the assertion that keeps it honest: the same token run twice gives 0 hits on the first pass — a cold bank cannot hit — and more than 0 on the second, because the same token routes to the same experts. 98 Swift tests, up from 97. The real-model hit rate is now historical, exactly as capital's digest is, because both were measured against machinery that has since changed; re-measuring is a second authorised run, not another fixture sources/DatacenterEngine/Qwen3_5Forward.swift, tests/DatacenterEngineTests/Qwen3_5MoEForwardTests.swift, docs/m1-gate.md, tracker. 98 Swift tests, 144 Python
2026-09-16 DC-087 resolved, and it changes the Metal plan rather than fixing a kernel. The divergence is not an indexing bug and not the partly filled group I guessed two rounds ago. Of twenty-four shapes one disagrees, and within it exactly one value of 128: index 117, where the GPU returns 0.0. That value's scale is a denormal — the CPU computes (5 - 41) * scale with scale ≈ -7e-39, below the fp32 normal floor, giving 2.6e-37, and Metal flushes the denormal operand to zero. Everything normal is identical. So Metal cannot reproduce the fp32 contract bit-for-bit wherever an operand is denormal, which is a property of the hardware and applies to every kernel, not to the unpack alone. That is exactly what I1/I2 care about, and it is now a decision rather than a surprise: either the contract flushes denormals on the CPU too — a renegotiation needing its own measurement and re-validation against the oracle — or the GPU is used only where denormals cannot arise. The kernel stays uncalled, the task's Done-when is rewritten to the achievable claim rather than quietly dropped, and the diagnostic that found this stays in the suite because "which shapes disagree" will be the first question about the next kernel too tests/DatacenterEngineTests/MetalUnpackTests.swift, DC-087, docs/m1-decisions.md
2026-09-16 The real-model run happened, and it produced the session's strongest result. With the operator's approval and the instruction to test that it works rather than for benchmark statistics — every node in the farm is doing coding work — all five prompts of the frozen set were run through the engine's prefill trace: capital 5 tokens, arithmetic 33, code 40, repeat 60, long 67, each exit 0 with a trace written, in 33/85/109/119/138 s. Those seconds are wall-clock under load and are recorded as existing, not as a baseline: capital took 45.6 s in the original run and 33 s here, and that is not a speedup claim because the earlier figure came from an idle machine. I1 was then verified directly on the real checkpoint rather than argued: capital and long re-run produce byte-identical traces, data.bin and manifest.json alike. And the strongest result: capital's digest is b8c976c5e7ba8816…, exactly the value recorded in the first gate run, whose trace was bit-identical to the contract — so the engine still emits the contract's exact bytes after the one-expert row read, the vectorised matmul, the eight-wide unpack and five language features. Each was proven bit-identical alone; this is the composition of all of them, on the real model, against the digest on record. Not measured: the contract half for the four new prompts, whose cost extrapolates to about five hours from 461.8 s per five-token prompt — so byte-identity remains proven on one prompt of five, and the other four are known to run rather than to match docs/m1-gate.md (all-five section), tracker. No code changed
2026-09-16 DC-015 done: the Swift language standard, measured rather than asserted. And then finished properly: MemberImportVisibility was first declared free on a swift build measurement (which does not compile the test targets), then declared expensive on an interrupted pass. Both were guesses. Iterating swift test under the flag and adding exactly the imports the compiler named took one round and one import — Qwen3_5ForwardTests.swift needed import DatacenterIR — and the suite then passes with the feature on, so it is now enforced like the rest. A feature that is not free is not the same as a feature that is expensive. Enforced: Package.swift now declares shardLanguageStandard and applies it to all six targets, so a target added later cannot quietly opt out: InferIsolatedConformances, ImmutableWeakCaptures, MemberImportVisibility and NonisolatedNonsendingByDefault. Each was probed by building with the flag and counting diagnostics — all four cost zero on a tree whose baseline is also zero, and all six targets build clean with them on. ExistentialAny is deliberately absent: it measures fourteen warnings, it is a mechanical pass across the engine's protocol types, and burying it in the same change would hide both. StrictConcurrency was probed too and is no longer separable — language mode 6 already enforces it. docs/swift-language-standard.md records the baseline, the four, the one left out with its count, the retired names, and the exact probe commands so the next person repeats the measurement instead of trusting it — with a note that CI cannot do this for you, because the macOS runner's Xcode is below the manifest's floor and the Swift job skips Package.swift, docs/swift-language-standard.md, tracker DC-015. 95 Swift tests green
2026-09-16 M1's gate status written down honestly, because the numbers in the doc were stale. docs/m1-gate.md now opens its status section by saying so: the surviving record of the first real-model run — the report JSON under .build/ is gone — reports 0.0191 tok/s / 52.2 s per step, and that was measured before three changes, each bit-identical and each measured locally: the int4 row read decoding one expert instead of a 256-expert stack, the matmul vectorised across outputs (2.98×) and the unpack vectorised and widened (1.83×). The arithmetic is why the first should dominate — ~2.0 B active parameters per token is ~9.3 s of unpacking at the old rate, and whole-stack decoding multiplied that by thirty-two to ~55 s, which is the measured 52.2 s/step reached independently. So "most of an order of magnitude better" is a prediction, and the doc says so in as many words. What is proven is unchanged and narrow: correctness IDENTICAL on one prompt of five, 83 tensors and 40 discrete decisions, digests equal; the cached path agreeing with the uncached one on tokens and margins; the 20 GB install building and verifying uncached at a 33.7 MB peak. What is not measured: the other four prompts, a re-measured baseline, the cache hit rate on the fixed engine, and the engine running against the install rather than the checkpoint. All four need one real-model run, which needs approval because the process peaks at 4.16 GB against ~4.5 GB usable. M1's gate is open, not passed — correctness holds on one prompt of five, and its throughput claim is retired pending that measurement docs/m1-gate.md (status section), tracker. No code changed
2026-09-16 DC-088 narrowed to one side of the wire: the writer is right. Recomputing every slab digest from data.bin in Python — slab by slab, for all eleven quantized fixture entries — matches the manifest exactly, so the format's definition of a slab digest is sound and self-consistent; the fault is in the Swift reader's transcription of it, which is where the earlier round found slab 0 disagreeing while slab 1 verified. The second probe I tried produced no evidence at all: swift test --filter matched zero tests, so nothing printed and I had left an instrumented branch in the source for nothing — reverted, and the correct filter recorded here (--filter InstallRowReadTests, not a single test name I had guessed at). Two rounds on this now, both short of the answer, and the honest statement is the narrow one: one side of the format is proven, the other is not, and the next step is a probe that actually runs tools/quantize.py (writer), DC-088. 95 Swift tests green, exit 0
2026-09-16 INCIDENT: this development node panicked twice, and the cause was my own memory budget, not a bug. While running the 35 B engine (~4.5 GB resident) and a 20 GB install build concurrently on an 8 GB machine, macOS grew swap to 13 swapfiles and LOW swap space; the system stopped responding for 90 s and the hardware watchdog panicked it — watchdog timeout: no checkins from watchdogd in 90 seconds, with the backtrace in AppleARMWatchdogTimer and kernel_task dominating the CPU. Both panics have the same signature and neither involves the engine's arithmetic. It is, however, the brief's own constraint made expensive: "usable RAM per node after macOS: assume ~4.5 GB. This is brutal; design for it." The design said so and the workflow ignored it. Standing changes, recorded in AGENTS.md so the next session inherits them: tiny-fixture work only (the same code paths at megabytes), one heavy job at a time, real-model runs only with explicit human approval and a swap check first, and .build/ excluded from Spotlight via .metadata_never_index — the 93 GB there was being re-indexed after every reboot, which is sustained I/O on a machine that had just panicked. Two useful facts survive the incident: the 4-bit install for the real model did build and complete (20 GB, peak footprint 4.44 GB, streaming writer, ~14 minutes), and the same build now streams to disk instead of accumulating twenty gigabytes in memory, which is what the first attempt did AGENTS.md (the node's hard limit), .build/.metadata_never_index, and the streaming InstallWriter in tools/quantize.py. No engine code changed
2026-09-16 The status that said the code does not run is gone, and the toolchain condition is now part of every claim. The operator asked for the documents to be fixed, and there were five, not the two I had named: README.md, AGENTS.md (the header and the trap entry), CONTRIBUTING.md, and two wiki pages — including .wiki/Home.md, whose front page still said there was "no source code, no build system and no release", the very extreme AGENTS.md had already recorded as one of the two wrong ones. Each now states what is measured and reproducible: on Swift 6.4 / Xcode 27 — the floor swift-tools-version:6.4 demands — swift build is clean and swift test --no-parallel gives 105 tests, 2 skipped, 0 failures, with 163 standard-library Python tests gated in CI on every push. The engine is incomplete: M0's dense path is bit-exact against the reference contract at fixture scale, M1 has run on the real 35B model with all five frozen prompts at exit 0 and two byte-identical re-runs, and M1's gate is open. The trap entry now records that the status has swung three times and states the rule that would have prevented all of them: a status claim must name the toolchain, because that is the whole difference between "does not build" (Xcode 26.x cannot parse the manifest) and "builds and passes" (Xcode 27). The header fix also removed a dangling sentence fragment — "The Swift engine under" had been left mid-clause when the status was spliced in README.md, AGENTS.md, CONTRIBUTING.md, .wiki/Home.md, .wiki/Roadmap.md, DC-102. 105 Swift tests, 163 Python
2026-09-16 I went looking for the next real gap in the engine and found coverage instead — so I am recording that rather than inventing work. The previous round's recommendation was to probe the Gated DeltaNet recurrence and the cache path. Both were already covered, and specifically: the recurrence by seven tests including testMultipleChunksMatchTheContract and testTheDecodeStepMatchesTheSequencePathOnTheLongCase, and the cache by five including testCachedGenerationProducesTheSameTokensAsUncached and testCachedAttentionIsBitIdenticalToTheSequenceAttention. The second of those is the exact property this audit was going to add — prefill plus decode against one long sequence, bit for bit — and it was already written and passing, which is the fifth time this session that checking first saved me from filing something false. The inventory now lives in docs/m1-gate.md with the test names, so the next round does not spend itself rediscovering it. What this leaves is a cleaner statement of the position: the engine is covered as far as a fixture can cover it, and every remaining measurement on M1's gate — the tok/s, the cache hit rate, the per-token bytes read, capital's digest — is a real-model measurement behind the same authorisation docs/m1-gate.md (covered already), DC-096. 100 Swift tests, 150 Python
2026-09-16 DC-088 done, on the seventh round, with three faults each proven fixed. Reading the region properly — instead of patching it from memory — found that InstallFile.rows carried two whole-entry digestMatches checks on the int4 path, one before the three payload reads and one after; every previous attempt changed exactly one, so the survivor rejected any tamper in the entry before a slab was consulted. That is why a test tampering the payload's last byte kept reporting the first expert as bad, and why I twice blamed my offset arithmetic. With both removed, a second fault became visible as a crash: the slab loop offset range-local buffers (codes holds only the rows just read) with a global slab index, so a read of expert 1 sliced codes[1024..<2048] on a 1024-byte buffer and trapped — which is also why the earlier instrument printed nothing before dying. The third fault was the guard, range.count % inner == 0 being 1 % 32 for one expert of a stack, so the branch was skipped and every read silently paid the whole-entry digest: correct, and exactly the cost this exists to remove. Proven now by two tests — a cold one-expert read cheaper than a third of the tensor (the assertion that fails when the guard stops matching), and a tampered last expert failing for itself while the first still reads. 96 Swift tests green under -Onone and -O. Two lessons worth more than the change: a guard that chooses between a fast path and a correct fallback needs a test that fails when the guard stops matching, and global indices and range-local offsets are different numbers — the leading-axis rule for the sixth time, in a new disguise sources/DatacenterEngine/Install.swift, tests/DatacenterEngineTests/InstallRowReadTests.swift, tools/quantize.py (writer, proven by its own test). 96 Swift tests, 137 Python
2026-09-16 D9: the fast matmul is bit-identical, and BLAS is measured out. orderedMatmul now vectorises across the output dimension with each lane accumulating over k ascending, so the rounding sequence is unchanged: bit-identical on every shape tried and the whole engine's 86 tests — golden and contract comparisons included — stay green. It buys 2.98× (4.1 → 6.8 GFLOP/s), and it is free precisely because the bits do not move. The decisive measurement is the other direction: Float.addingProduct differs from the ordered sum in 177 of 256 dot products, so an FMA — which is what every BLAS kernel uses internally — cannot be bit-identical, and cblas_sgemm is excluded for any op the gate compares. The same round measured where the time actually goes: dequantizeInt4 runs at 370.9 M values/s, so a 35 B token's ~3.45 B values cost ~9.3 s of unpacking, while its 6.9 GFLOP of matmul is only ~1.7 s at the scalar rate. Matmul is not the bottleneck. That also yields a prediction to test rather than believe: 40 layers of whole-stack decoding is ≈ 21.5 G values ≈ 58 s at the measured rate, which is the same number as the measured 52.2 s/step by a different route, so the int4 row-read fix made for memory should be worth most of an order of magnitude — unverified, and it needs a real-model run with the operator's approval sources/DatacenterEngine/Ops.swift (orderedMatmulVectorized, orderedMatmulScalar kept as the definition), tests/DatacenterEngineTests/OrderedMatmulTests.swift, docs/m1-decisions.md (D9). 86 Swift tests green under -Onone and -O
2026-09-16 DC-088's root cause: a guard condition, and the leading-axis trap for the fifth time. The slab branch tested range.count % inner == 0. For a stacked tensor a request for one expert is range.count = 1 while inner = 32 payload rows, so 1 % 32 is not zero, the branch was skipped, and every row read silently fell back to hashing the whole entry — correct, and precisely the 537 MB this work exists to remove. The slab index is the leading-axis index; the guard should have been range.lowerBound + range.count <= slabs.count, with the payload rows of slab i being i * inner ..< (i + 1) * inner. That fix was written and observed working: with it, the cold one-expert read fell below a third of the tensor, which is the measurement that proves the slab path actually runs rather than merely compiling — the same test that would have failed under the old guard. It did not land, because the accompanying tamper test failed on its own offset arithmetic (it tampered at offset + inner * codesPerRow + 1 and the untouched expert failed) and I ran out of room to find out where. So the tree was restored to the committed state, which is green at 95 tests, rather than left red for the sake of a nearly-finished change. Two things to carry forward: the guard bug is identified and its fix is measured, and the tamper test is a test bug, not a format bug DC-088, tools/quantize.py (writer, tested), sources/DatacenterEngine/Install.swift (reader, reverted). 95 Swift tests green, exit 0
2026-09-16 After withdrawing the false claim, I checked its real question properly — and it answers well. The confusion was real even though the claim was not: the real checkpoint nests its geometry under text_config, the fixture keeps it at the top level, and my script assumed one shape. So the question worth asking was never "does the fixture carry the field" but "does the engine, on the real model, apply the factor" — and the answer is in the artifact: .build/m1-install/install.json carries spec.config.partialRotaryFactor = 0.25 and spec.config.ropeTheta = 10000000. The validation model runs partial RoPE, the spec carries it, and the same file also carries family, passes, policy_files, source and schema — which is invariant I6's provenance, present in the artifact rather than promised in a document. Two rounds of my own error ended in a positive measurement about the engine, which is the honest shape of this session: the engine keeps being right, and my instruments keep being what needs checking .build/m1-install/install.json, DC-093 (withdrawn, then answered), docs/m1-gate.md. 97 Swift tests, 144 Python
2026-09-16 The D12 argument became an instrument, and the compiler corrected my first attempt. The bank size has been an open question for six rounds; rather than pick a number, the honest answer is a measurement — how often does routing actually repeat, at what RAM cost — so the run can sweep it. SHARD_EXPERT_SLOTS=<n> sets the bank for one process. My first version made it a static var, and Swift 6 refused to compile it as nonisolated mutable global state — which is exactly what the existing comment on that line had already predicted: "concurrency rejects mutable global state, and it should become a field of the IR's policy when there is a measurement to put in it". So it is a static let, read once, which is what a sweep wants anyway (one process per setting, since two sizes cannot share banks). The parse is a separate function, tested against "0", "-3", "many" and "" — a bank of zero would silently disable the cache the measurement exists to observe. Also corrected this round: I had been listing DC-090 (the sign of zero) as a blocker on M1's gate. It is not — Metal is not on the engine's expert path at all; MetalUnpack is unused, so the divergence only affects whether that optimisation could ever be adopted. Two real blockers remain: the run, and the size it should make decidable sources/DatacenterEngine/Qwen3_5Forward.swift, tests/DatacenterEngineTests/SlotBudgetTests.swift, docs/m1-gate.md, DC-092. 101 Swift tests, 156 Python
2026-09-16 M0 was finished all along, and my status rewrite said it was not. Asked what was next for M0, I read docs/m0-gate.md instead of answering from memory — and it says status: passing, with a recorded run on Qwen/Qwen3.5-2B at revision 15852e8c…, three frozen prompts, 40,683,520 bytes of trace data identical to the Python contract and every discrete decision matching the reference implementation, on one Mac mini M2 with 8 GB. Every M0 task is Done: the trace harness (DC-024), the diff harness (DC-025), the reference contracts (DC-027), M0a's proven-able-to-fail harness (DC-028), M0b's bf16 layer-by-layer match (DC-029), M0c's cost of 4-bit (DC-031), and the gate itself (DC-026). The answer is that nothing is left in M0 — so the round's work became the correction: six documents I had rewritten hours earlier calling M0 "proven at fixture scale, not by running a dense model on this node". I wrote that from the fixture suites I had spent the session inside and never checked it against the gate doc — the same failure as the invented geometry and the JSON lookup in the wrong nesting: the repository held the answer and I did not read it. The corrected paragraph also states two things the earlier wording hid: the pinned model is 2B where the brief says "~1B", and the oracle half was renegotiated (bit-matching torch is impossible, D3/R13), so against the reference the gate asserts exact decisions plus recorded closeness rather than bytes — a recorded weakening of the brief's literal words, which is what rule 3 asks for README.md, AGENTS.md, CONTRIBUTING.md, .wiki/Home.md, .wiki/Roadmap.md, docs/brief-response.md, DC-104. 105 Swift tests, 174 Python
2026-09-16 The fixture omitted the one config field the brief warns about, and comparing it against the checkpoint found it. Last round's lesson was that I had invented geometry the repository already had, so this round audited the engine's numbers against the checkpoint instead of against my memory — and the checkpoint's config.json makes it easy: partial_rotary_factor is 0.25, rope_theta 1e7, full_attention_interval 4, head_dim 256. The engine is fine: Qwen3_5Forward takes config.partialRotaryFactor ?? 1.0, rotates only the first rotary channels with applyPartialRope, and both importers read rope_parameters.partial_rotary_factor; the op-level contract vectors test it at 0.5 and the qwen3.5 fixture at 0.5 too. But Fixtures/tiny-qwen36/config.json omits the field — as it also omits rope_theta and layer_types — so ?? 1.0 applies and the MoE forward that M1's gate rests on has never run with a partial rotary factor, while the model it gates on always does. Nothing is known to be broken, and that is the point: the two sides share the default, so the fixture cannot catch a mistake in it, and the brief names RoPE application order as a classic failure that looks fine and is wrong. Filed as DC-093 and listed under not measured in the gate doc, because "we agree on a default" is not the same claim as "we agree with the checkpoint" docs/m1-gate.md (not-measured), DC-093, tools/make_contract_vectors.py:118, tests/.../Fixtures/tiny-qwen36/config.json. 97 Swift tests, 144 Python
2026-09-16 DC-087's residue is not denormals, and the shape of the failure names the cause. With D11's operand flush in place, three rounds of hunting the remaining divergence produced a clean elimination instead of a fix: the full grid fails on 6 shapes, each with only 1–4 differing values out of hundreds, at scattered indices — and flushing the product as well as the operand changed nothing at all, which rules denormals out. What the failures do share is rows: columns = 4, group = 1, rows = 1 passes while rows = 3 fails, and columns = 128, group = 1 fails at rows = 1 only at index 117 of 128 — so this is a row-index or row-stride difference in the shader, not arithmetic. That is a much better place to start than "the GPU disagrees". The result-flush experiment was reverted: it did not fix anything, so it did not earn a second change to the contract — a rule worth keeping, since D11 was decided by measurement and this would have been decided by hope. The tree is green at 97 Swift tests and 139 Python, and D11's own claim — the operand flush, chosen by the operator — is untouched and proven sources/DatacenterEngine/Install.swift, tools/quantize.py (both reverted to D11 as decided), DC-087. 97 Swift tests, 139 Python
2026-09-16 DC-088: the instrument ran this time, and it eliminated the arithmetic. Written as a real file rather than through a heredoc (the last one printed its own format string because of multi-line interpolation in a shell heredoc — a tooling failure that cost a round). It prints the numbers the slab path derives and asserts nothing, and the result is that all of them are consistent with the manifest: for the stacked expert tensors leading = 8, rows = 256, inner = 32, slabs = 8, codesPerRow = 32 — the padded width, which is 64 here although the tensor has 16 columns, so the padding is stored and the reader is right to use it — groupsPerRow = 1, and payload == nbytes for all eleven quantized entries. Slab 0 of down_proj is therefore codes 1024 of 8192, scales 128 of 1024, zeros 32 of 256, all in range. So the remaining fault is inside the slab branch's slicing, before the print that never appeared — the instrument from the previous round trapped with signal 5 and printed nothing, which places it after the guard and before the first hash. That is a real narrowing and it is where the next session's first print belongs: one line per slice, to see which of the three traps. Recorded, not attempted: six rounds have gone to this task and a seventh fragment would not be the disciplined move DC-088, tools/quantize.py (writer, tested and proven), sources/DatacenterEngine/Install.swift (reader, reverted). 95 Swift tests green, exit 0
2026-09-16 D12's constraint is a test now, so the number is visible without reading the wiki. The audit that found the 14.5 GB mistake also found that nothing checks the literal expertSlotsPerLayer = 16 — no fixture can, because a fixture's experts are kilobytes. SlotBudgetTests does the arithmetic the fixture cannot, on the checkpoint's geometry (2048, 512, 40 layers, from config.json and confirmed in the real spec.config), and asserts the total against the brief's own limit of ~4.5 GB usable per node. It fails today and XCTest prints why: 8053063680 against 4500000000 — 8053 MB against 4500 MB. It is marked XCTExpectFailure for three reasons that matter more than the green tick: anybody running swift test sees D12 without opening a wiki page; the suite stays honestly green, because a known-unmet constraint recorded as met would be a lie; and when the bank is finally sized from a budget the test reports an unexpected pass, which forces the marker to be deleted deliberately rather than left behind. Alongside it, two tests pin the arithmetic itself (12,582,912 bytes per expert; 8,053,063,680 at 16 slots) and one records what a 1.5 GB budget allows — two whole slots per layer — so the pending decision has its answer attached tests/DatacenterEngineTests/SlotBudgetTests.swift, docs/m1-decisions.md (D12), DC-092. 100 Swift tests, 2 skipped, 0 failures
2026-09-16 I4 and L2 audited, and the audit's first answer was wrong in an instructive way. Checking "policy is data, not code" against the artifact found every one of the real install's 27 roles has a policy entry, and that the policy carries a why block justifying each role against an invariant — routers bf16 because I3's discrete decisions must survive, the routed experts because they are "92.9 % of this model's parameters", the shared expert because its error is not amortised when it is active on every token, and the Gated DeltaNet's decay because it is exponentiated, where "a 4-bit error is not a small error". The missing-role behaviour that AGENTS.md states as a trap is tested four times. Then the reverse check flagged mlp.down, mlp.gate and mlp.up as entries no tensor uses — and before writing "dead entries" into the tracker I looked, and they are the dense fixture's roles: the policy is shared across families by I4's own design, and the union is what must match. So the check became a test of the union, which fails on rot in either direction. Importers measure 144–235 lines against L2's 500–800 ceiling; sharding policy is absent because the brief puts it in M2. Third surface audited this way, third time the work was already done — and second time my instrument's first reading would have produced a false finding had I not checked it tools/test_quantize.py, docs/m1-gate.md, DC-099. 100 Swift tests, 156 Python
2026-09-16 The unpack kernel: 1.60×, and it was the easier one. D9 said the frame goes to the unpack (370.9 M values/s) rather than the matmul (4.1 GFLOP/s), and the unpack turned out to be easy for a structural reason worth remembering: its only floating-point operation is one multiply, Float(code - zero) * scale, with no summation anywhere — so there is no accumulation order to preserve and a vector formulation is bit-identical by construction, not by luck. That is the opposite of the GEMM, where FMA makes BLAS unusable. Measured: 648.3 → 1038.3 M values/s (1.60×), so one 35 B token's 3.45 G values take 3.3 s instead of 5.3 s. The win is not the vector ALU but the group: the scale and zero point are per sixty-four values and the scalar loop reloaded both per element. Bit-identical on a grid of sixty shapes (columns not a multiple of four, the padded tail, a group of one, one row) and on the end-to-end golden tests. Two honest notes: the obvious next step is eight nibbles per iteration with integer lanes, and it is not measured; and the grid cost an hour to a Swift footgun — % keeps the dividend's sign, so (-5) % 4 == -1 and the test compared the implementations on an impossible layout sources/DatacenterEngine/Install.swift (int4Layout, vector dequantizeInt4, dequantizeInt4Scalar kept as the definition, and an error message that now says which condition failed), tests/DatacenterEngineTests/Int4UnpackTests.swift, docs/m1-decisions.md. 90 Swift tests green under -Onone and -O
2026-09-16 I6 is half-implemented, and the half that is missing is the half that makes provenance worth having. Auditing the brief's own words against the artifact — the method that found the alignment gap the round before — showed the install cannot answer "which checkpoint is this". source.files is {} in the real install and in both committed fixtures, so "sha256 of each source weight file" is unmet, and revision is always "local" while repo holds either a commit hash or a fixture name, one field carrying two meanings. The other half is genuinely there: passes, policy_files (relative in the fixtures, so nothing platform-specific is committed), 693 per-tensor digests of the converted payloads, the spec and the family. I checked for a privacy problem before reporting one, and there is none: git grep "/Users/" over tracked files returns nothing, and the absolute path that does exist is in a gitignored .build/ artifact. Filed as DC-098/D14 with the remedy and its measured cost — the uncached digest path that verified 20 GB in ~20 s at a 34 MB peak does the same job on the source files, data.bin does not change, so nothing needs rebuilding and the work batches with the next real-model run instead of being a reason for one docs/m1-decisions.md (D14), DC-098, .build/m1-install/install.json, tests/.../Fixtures/tiny-qwen3{5,6}/install/install.json. 100 Swift tests, 150 Python
2026-09-16 The 5 GB disk floor is enforced, not promised. After the two watchdog panics, the fix is mechanical rather than aspirational: tools/disk_watchdog.py polls free space and, below the floor, writes .build/DISK_STOP before doing anything else and then terminates the heavy jobs that are running — the engine, the contract, the install builder — with SIGTERM first and SIGKILL after a grace period. It matches tools by name rather than killing every Python, because the agent harness is a Python process too, and it never terminates itself. The other half is tools/check_disk_headroom.py: quantize.py, the contract CLIs and run_m1_gate.py call require_headroom() before touching a model, so nothing heavy starts below the floor either, and a stop marker refuses a new run until an operator clears it deliberately — a run that tripped the limit must not resume by itself. The marker is written first precisely because the window between stopping the old jobs and the next poll is when a new one would otherwise slip in. Four tests cover the refusals (including that the guard does not clear the marker it refuses on), and they are stdlib-only so CI runs them. The watchdog's first version was too twitchy and it proved it on itself: it acted on a single reading of 4.67 GB that was 17.55 GB six seconds later — APFS "purgeable" space appears and disappears as the system reclaims caches — and killed a read-only install verification that had done nothing wrong. It now requires three readings in a row (--consecutive 3) and writes the marker before stopping the jobs, so a job starting in that window still refuses to run; four more tests pin the decision, with the reading faked and the signal replaced by a recorder, and one asserts the watchdog never matches itself or the agent harness. Measured state when it went live: 18 GB free, the watchdog running at pid 1352, .build holding 67 GB of checkpoint and 20 GB of install tools/disk_watchdog.py, tools/check_disk_headroom.py, tools/test_disk_headroom.py, and the floor written into AGENTS.md where the next session reads it before running anything. 129 Python tests
2026-09-16 Per-slab digests: the writer half landed, the reader half was reverted rather than shipped. The cost being attacked is real: InstallFile.rows verifies an entry by hashing its whole payload, so handing back one expert of a [256, 1024, 2048] stack reads 537 MB first — across the real model's eighty expert tensors that is tens of gigabytes of extra I/O per process (an estimate from the entry sizes, not yet measured end to end). The install format now carries slab_sha256, one digest per leading-axis slab, computed over that slab's three section-major ranges in a fixed order — additive and forward-compatible, so a schema-1 reader ignores it and the fixture regenerates byte-identically as far as values go. The reader side was written, and it disagreed with the writer: reading an untouched expert (slab 0) threw a digest mismatch while slab 1 verified, which means the two implementations do not define slab 0's digest the same way. I did not diagnose it before running out of room, so it was reverted and the tree left green rather than committed with an unexplained disagreement — which is the rule this project keeps re-learning, and the second time in this session that a plausible fast path was withdrawn on evidence. The repro is preserved as the test that found it and filed as DC-088 tools/quantize.py (InstallWriter.slab_digests), sources/DatacenterEngine/Install.swift (slab_sha256 present and deliberately unused), tests/.../Fixtures/tiny-qwen36/install/ regenerated, DC-088. 95 Swift tests green
2026-09-16 DC-086 is Done, with the measurement it asked for. Both payload sources now read uncached — UncachedFile in Swift for the install and for the checkpoint's streaming row path, the same in Python for the tooling — and verification moved from hashing the whole payload on open to checking each payload the first time it is read, in both languages, with I6 preserved because a tampered payload still cannot be decoded into plausible weights. The measurement: the full 20 GB install verifies on this 8 GB node in about twenty seconds, peak memory footprint 33.7 MB, free disk steady at 16 GB. Two numbers make the comparison: the previous attempt at the same verification took free disk from 17 GB to 2.96 GB (the page cache became memory pressure became swap) and the previous verifier held the whole file, so it would have needed ~20 GB resident. Five new Swift tests cover the streaming read (every row-addressable tensor reads the same both ways, rank-1 falls back, out-of-range still throws, the expert provider streams and never touches the mapped path) and two new Python tests cover the uncached primitives and the moved verification sources/DatacenterEngine/UncachedFile.swift, Safetensors.swift (rowsStreaming + one shared decoder), Install.swift, ExpertProvider.swift; tools/quantize.py (open_uncached, pread_exact, digest_of); tests/DatacenterEngineTests/StreamingReadTests.swift, UncachedFileTests.swift. 77 Swift tests, 136 Python
2026-09-16 DC-098 is closed, and the item I kept calling open was never open. The audit found I6 half-implemented — source.files empty, revision a placeholder — and over four rounds both halves were fixed and pinned: source digests verified against an independently computed sha256 and against the committed artifacts, the placeholder now impossible by a property test, and --repo/--revision inputs for what a build cannot know. But I also wrote, twice, that the fixture builders would drop the digests on regeneration, and that was false. They call quantize.build_install(FIXTURE, …) — the function that now digests — and the check that settles it is one line: for both fixtures, the committed manifest equals digest_snapshot of the fixture directory, True and True. The false claim came from a regex of mine that matched no source = { block, and I turned my tool's failure to match into a statement about the repository. That is the third time an instrument of mine produced a false reading — a heredoc, a JSON lookup in the wrong nesting, and now a regex — and each time the repository was right. DC-098 is Done with all three items verified tools/quantize.py:634, tools/make_tiny_qwen3{5,6}_checkpoint.py, DC-098. 100 Swift tests, 156 Python
2026-09-16 A brief requirement turned out to be unimplemented and unnecessary, and measuring it is what made that a decision rather than a defect. L3 asks for experts "repacked contiguous as [gate|up|down], 16 KB aligned"; nothing in the repository mentioned 16 KB at all. Rather than implement it, or ignore it, the question was why the brief asks: O_DIRECT is what needs 16 KB, because it requires the buffer, the length and the file offset aligned to the device block size — and this engine reads through F_NOCACHE + pread, which needs none of that. O_DIRECT appears exactly once in the whole repository, in a comment quoting the brief. The measurement, on the real 20 GB install's 693 tensors: 0 offsets fail 2-byte or 64-byte alignment, 375 fail 512, 667 fail 4096, 688 fail 16384 — so the layout is deliberately 64-byte aligned, which is what a SIMD4<Float> load and a pread want, and the reader uses explicit loadUnaligned everywhere. Filed as D13 with the revisit condition stated as a single measurement: if the read path ever moves to O_DIRECT, the install is rewritten and every trace regenerated, because the offsets change. That is the difference between deferring a requirement and forgetting one docs/m1-decisions.md (D13), DC-097, .build/m1-install/install.json. 100 Swift tests, 150 Python
2026-09-15 The 4-bit path is bit-reproducible end to end (DC-031 closed). The engine now reads the install: Swift decodes the same packed codes, group scales and zero points as the Python pass, and the two agree byte for byte on all three frozen prompts — 40,683,520 bytes of trace data, matching digests, the differ reporting IDENTICAL. Unlike the bf16 path this comparison is byte equality rather than a tolerance, and legitimately so: the codes, zero points and scales are integers and exact arithmetic, so there is no rounding to disagree about, and a difference would mean a wrong layout rather than a different order. Three things came with it: the install carries its own IR spec so the engine needs nothing beside it and the forward pass's configuration is reconstructed from that spec (which is the check that L1's spec is sufficient to run a model rather than a description of one); the loader detects an install by its header, so one binary runs a checkpoint or an install; and the tiny fixture generator now writes a tiny install plus the contract's output on it, so the int4 path is covered by CI rather than only by a run that needs a 4.5 GB checkpoint. The measured cost stands as recorded in docs/m0c-quantization.md: 2.51× smaller, 7.71 bits/weight, 28 of 29 decisions preserved against the fp32 reference sources/DatacenterEngine/Install.swift, tests/DatacenterEngineTests/Qwen3_5ForwardTests.swift (3 new tests, including one that re-derives the format's arithmetic longhand so a wrong nibble order cannot pass), and the record in docs/m0c-quantization.md
2026-09-15 DC-032 closed: the experts are read by index, not as a stack. A layer's experts are 805 M parameters — 3.2 GB in fp32, 1.6 GB in bf16, against about 4.5 GB of usable memory per node — so the kernel now asks an ExpertWeightProvider for the experts the router chose and never sees the stack. The layout makes this cheap: the checkpoint's leading axis is the expert, so one expert is exactly one row of the stacked tensor and a fetch is a single row range. StackedExpertProvider reads it, ExpertSlotCache bounds what stays resident and counts, CountingExpertProvider measures. Two decisions are written down rather than implied: the kernel asks in ascending expert index, which is the contract's accumulation order and D4's ring order — the read order is the reduction order, and a cache that reordered reads for the disk's benefit would change the arithmetic; and the streaming path and the array path go through one kernel implementation, so they cannot drift. Measured on the tiny checkpoint: a nine-token prompt reads distinct chosen experts × 2 slices per layer and nothing else, and caches of 1, 2 and 8 slots give output bit-identical to the array path. A bug found on the way: I first read 2·intermediate rows per expert, treating the flat tensor as [experts·2·inter, hidden]; the reader reports the leading axis as the row count, so the correct unit is one row, and the provider now checks the width against the mixture's geometry rather than trusting the layout. What is not measured: the 67 GB checkpoint has not been fetched, so M1's throughput baseline and hit rate remain open (DC-034, now next) sources/DatacenterEngine/ExpertProvider.swift, the provider turn in MixtureOfExperts.swift, and tests/DatacenterEngineTests/ExpertProviderTests.swift (7 tests). 53 engine tests, green under -Onone and -O
2026-09-15 The Gated DeltaNet layer and this family's RoPE are transcribed and checked against the reference module (DC-029). Qwen3_5TextRotaryEmbedding turned out to hold two facts that would have been silently wrong if guessed: partial_rotary_factor is 0.25, so only the first 64 of 256 head dims rotate, and rope_theta is 1e7, not qwen3's 1e6. apply_rotary_pos_emb:674 is partial by construction — it rotates the first rotary_dim dims and passes the rest through — so a full-head rotation would have been wrong. The multimodal recomposition_frequencies machinery is a no-op for text-only input, and that is a fact about the caller (Qwen3_5TextModel.forward:1260 expands arange(seq) to four identical grids), recorded with its line numbers rather than as a conclusion. tools/ordered_qwen35.py then implements the whole GDN layer and 7 tests compare it against transformers' own Qwen3_5GatedDeltaNet at 20, 37 and 70 positions — the last spanning more than one chunk — plus conv causality, softplus's threshold and the sign of the decay. A latent bug was found and fixed on the way: the conv loop indexed batch 0 from inside the loop, so a batch of two would have returned the first sequence twice, and shapes hid it tools/ordered_qwen35.py, tools/test_ordered_qwen35.py, and the transcription in docs/reference-qwen35-2b.md; the whole family's forward is what remains before the Swift kernel
2026-09-15 DC-014 — the project's language standard is Swift 6.4 on Xcode 27 (swift-tools-version:6.4, Swift 6 language mode); the earlier 6.3.3 line is superseded AGENTS.md, CONTRIBUTING.md, the Testbed page and the Architecture stack row; the reference machine reports Apple Swift 6.4 / Xcode 27.0
2026-09-15 The GDN decode step exists and is verified against the sequence path — the substance of the cache, before the wiring. GatedDeltaNet.decodeStep carries a State of two pieces: the convolution's window (the last kernel - 1 raw projections per channel, which causal_conv1d_update:252 concatenates onto the new input, and which a zero-padded conv would silently replace with zeros where history belongs) and the recurrent state [heads, keyHeadDim, valueHeadDim]. It is checked against the layer's sequence path on the golden vectors — the 70-position case that spans chunks and the asymmetric two-key-heads-to-four case, because the grouped-query repetition is a step a symmetric fixture cannot check — at a tolerance rather than at the bit, which is D8's consequence stated as a test: the decode step and the sequence path are the same recurrence, and the sequence path is still bit-identical to the contract. Two lines in the transcription of torch_recurrent_gated_delta_rule:440 are worth naming because both produce plausible numbers when wrong: the axis convention, where a transposed rule broadcasts happily, and the unconditional query / sqrt(head_dim) after the l2norm, whose absence leaves the output out by a constant factor. What is not yet wired: a prefill that leaves the states behind, the attention layers' KV cache, and the cached generate loop — deliberately after the unit, because the wiring is where a state goes into the wrong layer or the wrong batch element and still produces plausible text sources/DatacenterEngine/GatedDeltaNet.swift (State and decodeStep), two tests in GatedDeltaNetTests.swift, and the note in docs/m1-decisions.md. 57 engine tests, green under -Onone and -O; 125 Python tests
2026-09-15 Swift has a CI job. .github/workflows/swift.yml builds and tests the package on a clean macOS runner and prints the toolchain first, because the manifest asks for swift-tools-version:6.4 and the runner's version is the one thing that can silently disagree. Also found while writing the manifest: SwiftPM resolves Sources//Tests/ case-insensitively on macOS, so the lower-case directories this repository uses (matching the sister project) would have failed on a case-sensitive filesystem or on a runner — the target paths are now declared explicitly The workflow itself, and docs/repository-layout.md under "Rules that came out of building it"
2026-09-15 M1's gate passes on the real model, and its throughput number is the finding. One command, all three parts: correctness IDENTICAL (83 tensors, 40 discrete decisions, digests equal), peak memory 4.16 GB, hit rate 0.0000 over 2240 requests, throughput 0.0191 tok/s = 52.2 s per token. That last figure is not a tuning result: it is what a faithful fp32 CPU port with no KV cache costs, because every generation step re-runs the whole sequence, so a step grows with context — at a hundred tokens this arithmetic is minutes per token. M0 deferred the cache on purpose ("a cache is a second numeric path through attention and M0's job is to establish one correct path before there are two") and that deferral has now been paid for honestly: for M1 the KV cache is not an optimisation, it comes before Metal, because a kernel is a constant factor and this is a growth term. Two fixes landed with it: the expert counters now report elements and bytes rather than a field named rowsRead that held elements in one counter and rows in another — the brief's currency is bytes from the SSD, and the unit was mislabelled in a way that still looked like a plausible number — and the gate takes --only, because the contract side costs about ten times the engine's wall time and a five-prompt run is hours rather than minutes. The traffic behind one token: 2240 requests, 3.52 G elements, 7.05 GB read from the SSD in bf16 and 14.1 GB held in fp32 docs/m1-gate.md (the gate table and what the throughput means), tools/run_m1_gate.py (--only), and the units in sources/DatacenterEngine/ExpertProvider.swift and sources/DatacenterTrace/main.swift. 121 Python tests, 55 engine tests. Next: the remaining four prompts as a background run, then the KV cache
2026-09-15 D1–D7 decided, and DC-002 closed with them. The operator chose the M0 model (dense Qwen3.5-2B at 4-bit), kept Qwen3.6-35B-A3B as M1, restricted the reference to the four M2 nodes, and settled the name as TinyTitan Datacenter; the three delegated decisions were taken on evidence — fp32 compute on both sides for the gate, per-expert exchange with ascending expert-id accumulation, fp32 router logits with a documented tie-break docs/m0-decisions.md; the README now names the project and lists the three verified target models, so all four original README findings are fixed
2026-09-15 The M0 gate passes on the pinned model (DC-026 closed). The prompt set is frozen in tools/m0_prompts.json (sha256 3455d25e…) with token ids resolved once and committed, so the gate does not depend on a tokenizer version. On Qwen/Qwen3.5-2B at revision 15852e8c…, three prompts (5, 18 and 6 tokens): 40,683,520 bytes of trace data identical between the engine and the contract on every prompt, and every discrete decision matching the reference implementation. Worst relative difference against the oracle: 2.0e-06 on each. The interesting column is the smallest top-1 margin, which varies by two orders of magnitude between prompts — 1.1554, 0.0862, 0.0332 — while the divergence stays flat: the safety of a prompt is a property of the prompt, not of the implementation, and no tolerance on the numbers can tell you which prompt you are looking at. That is the argument for asserting decisions separately, now with three measurements behind it instead of one. tools/run_m0_gate.py reruns the whole thing and writes a machine-readable report; the record is docs/m0-gate.md, which also states what the gate deliberately does not cover (long context, 4-bit, Metal) and why tools/run_m0_gate.py, tools/m0_prompts.json, docs/m0-gate.md
2026-09-15 DC-003 — repository scaffolding: .gitignore, AGENTS.md, CONTRIBUTING.md, SECURITY.md, issue and pull-request templates, twelve repository topics and the seven phase milestones on GitHub The files are on main; the milestones list as #1–#7 in the GitHub API, each named after a phase
2026-09-15 The 4-bit path reaches the mixture, and the router's decisions survive it (DC-034 started). The quantizer and both install readers assumed rank 2, so the stacked expert tensors were refused outright; they now treat the leading axis as the rows at any rank, which is the same rule the streaming provider uses and the only one that keeps a quantisation group inside one expert. Measured on the tiny qwen3_5_moe checkpoint: 18 of 18 router top-k rows identical to the fp32 run — the policy keeps router.logits at bf16 for exactly this reason, and this is the check rather than the assertion — while the logits move by ~2.0e-01 relative, the same order M0c measured on the 2 B model. Two decisions are recorded in tools/quant_policy.json: the routed experts go to int4 because they are 92.9% of the parameters, and the shared expert stays at bf16 because it is active on every token where a routed expert serves eight in 256, so its error is not amortised and quantizing it would spend accuracy on the dense path to save nothing measurable. Three bugs found on the way, each invisible until the path was run: mixer_weights called layer_weights with the wrong arity, so the MoE streaming contract had never executed; the contract read numExpertsPerTok where the IR spells it numExpertsPerToken, which left top_k None and only failed against a spec the engine had emitted — the L1 design catching its own drift; and the Python install reader sized its code block from shape[0], the same mistake as the Swift reader, which the new rank-3 round-trip test now guards. The 35 GB download is running in the background tools/quant_policy.json (30 roles), tools/quantize.py (flat_rows), tools/test_ordered_qwen36_quant.py (3 tests), and the rank-generalised readers in quantize.py and Install.swift. 114 Python tests, 53 engine tests
2026-09-15 The Gated DeltaNet is transcribed from the reference, and the transcription is checked against the reference's own function (DC-029 under way). Every step of Qwen3_5GatedDeltaNet.forward:550 and torch_chunk_gated_delta_rule:301 is now recorded in docs/reference-qwen35-2b.md with line numbers: the depthwise conv (kernel 4, no bias, left padding 3, silu, truncate), beta = sigmoid(b) in the model dtype, g = −exp(A_log)·softplus(a + dt_bias) in fp32, the chunked rule itself (padding to 64, cum_decay, the -inf mask applied before the exp, the UT transform, the unit-lower-triangular solve, the sequential chunk scan), and the gated norm whose whole point is two orderings — the normalised value is rounded to bf16 before the weight multiply and the gate is activated in fp32 after it. tools/ordered_gdn.py implements it in the contract's order and six tests compare it against transformers' actual torch_chunk_gated_delta_rule across one chunk, several chunks, a small chunk size, an initial state and the returned final state — the tests pass. One trap caught by reading rather than by running: the attention's output gate is split per head ([tokens, heads, 2·head_dim] halved along the last axis), not into all queries then all gates; the global reading produces identical shapes and a wrong model. Also fixed: ordered_matmul only handled 2-D operands, which the delta rule's per-head contractions exposed The transcription is in docs/reference-qwen35-2b.md; the implementation and its 10 tests are in tools/ordered_gdn.py and tools/test_ordered_gdn.py. The dense golden vectors are byte-identical after the ordered_matmul generalisation, which is the check that the fix did not disturb the qwen3 contract
2026-09-15 M1's text tower has a contract, and building it found two silent bugs in the Gated DeltaNet. tools/ordered_qwen36.py composes the mixture behind this family's norms and residual and reuses ordered_qwen35 for the attention and the Gated DeltaNet — reuse that is licensed by compare_reference_modules.py having proved those entities identical, not by assumption. tools/test_ordered_qwen36.py then checks the composition against the reference's own Qwen3_5MoeTextModel, layer by layer, and checks the router's decisions at every layer as its own assertion (I3). The bugs it found are one asymmetry: this family has 16 key heads to 32 value heads where the 2 B model has 16 and 16, so (1) the IR derived the key head count from the value head count — the contract and Qwen3_5Forward.swift:147 both computed a key head width of 64 instead of 128, wrong in every Gated DeltaNet layer with no shape error anywhere — and the IR now carries linearKeyHeads as its own field; and (2) the grouped-query head expansion (query.repeat_interleave(num_v_heads // num_k_heads, dim=2)) was missing entirely, so the delta rule paired the wrong heads. The contract is fixed and its tests pass; the Swift kernel is not, and is filed as DC-038 rather than patched quietly. Both were caught by a tiny fixture built with the real model's asymmetric head counts instead of a convenient symmetric one — the argument for taking fixture numbers from the model rather than from what is easy. Also recorded: full_attention_interval is consumed by the reference's constructor and is not an attribute afterwards, so the checkpoint's config.json is the only place to read it tools/ordered_qwen36.py, tools/test_ordered_qwen36.py (7 tests), the linearKeyHeads field in the IR, and the two bugs written up in docs/reference-qwen36-35b-a3b.md. One of the two failures in the first green run was my own harness comparing every token against the reference's last row
2026-09-15 M0b is met: the engine and the contract are bit-identical on the real 2 B model (DC-029 closed). Qwen/Qwen3.5-2B at revision 15852e8c…, 8 tokens: 11,223,040 bytes of data.bin identical, 51 of 51 tensors, digest a1503f64…, the project's own differ reporting IDENTICAL and exit 0. Neither implementation can hold the model — 2 B parameters are 8 GB in fp32 against ~4.5 GB usable — so both stream: each reads a layer's tensors, uses them and releases them, and each reads the embedding a row at a time and the tied head in blocks of vocabulary rows. The symmetry is the point; a bit-exactness target that cannot run where the engine runs is not a target. Two structural pieces came with it: the tensor names now travel as data (datacenter-trace --emit-spec writes the engine's IRSpec, and the Python side consumes it rather than carrying a second copy of the importer's table, so L2 holds across languages and L1's spec file is finally an artifact something else reads), and tools/check_engine_contract.py reads the checkpoint's model_type and picks the contract itself. The two Python paths — resident and streaming — are checked against each other on the committed tiny checkpoint, so the path that otherwise only runs where a 5 GB file exists is covered by CI. Python 72.6 s for 8 tokens against the engine's 18.2 s: the contract is the slow side and does not need to be otherwise tools/check_engine_contract.py reruns the whole claim for either family; tools/ordered_qwen35_trace.py is the streaming contract, and docs/trace-format.md carries the numbers. DC-026, the M0 gate itself, is now a frozen prompt set and a full run away
2026-09-15 DC-019 — baselines measured and the target models verified: 4 KB round trip 645 µs TCP / 687 µs UDP against a 515 µs empty ping (latency-bound, not bandwidth-bound); internal SSD 1161 MB/s sequential but 108 MB/s random 16 KB at QD1; ~2.5 GB reclaimable memory on an idle node; no external NVMe attached; all three target configs fetched from their own checkpoints with revision hashes Testbed's "Measured baselines" and the Target models page. Each probe was built to fail loudly — the first attempts produced 13 µs round trips (a dead echo server) and 13 GB/s reads (a file that fit in cache), and both were discarded rather than reported
2026-09-15 DC-001 — the wiki was bootstrapped with Home, Roadmap, Project Tracker, Architecture, Testbed and Glossary, plus sidebar and footer navigation The wiki repository's own history; every internal link resolves
2026-09-15 README fixes — the stale NVMAI/TinyTitan reference (finding 2), the logo that was both committed and hosted from user-attachments (finding 4: the README now renders the committed file), and the Thunderbold typo The README diff of the same date; the remainder of DC-002 is listed above under P0
2026-09-15 The 35 B checkpoint is complete and the importer has met its full inventory. 26 of 26 shards, 1045 tensors in the index, and --emit-spec through the sharded loader returns 693 text tensors — which is exactly the count derived from the shard headers before a single weight was downloaded (1045 − 333 vision − 19 MTP). Two independent routes to the same number: the headers said what the file contains, and the importer now maps it. That closes the loop on the sharding work of the last two rounds, and it means the importer is no longer tested only against a tiny fixture. Next: the first real M1 gate run — one prompt, engine against contract, with wall time, peak memory and the router's decisions on the actual model .build/spec-36b.json (693 tensors, family qwen3_5_moe), emitted in 0.04 s. The gate run itself is the next round
2026-09-15 D8 decided: the chunked Gated DeltaNet rule is authoritative, and the cache is a second numeric path. With the reference's two paths measured against each other (ours agree to 3.2e-07; the reference's own two disagree by 3.4, a complete divergence), the question went to the engineer rather than being guessed past, and the answer is the chunked rule — the one prefill uses and the one M0 verified. What follows is written down in docs/m1-decisions.md: M1's bit-identity claim stays where it was measured (the uncached path against the contract); the cached path is validated against the uncached engine at the agreed protocol with the discrete decisions asserted exactly; and I1 is not weakened, because the same prompt in the same mode remains deterministic and what I1 never promised is that two different numeric paths agree. The discrepancy in the reference is recorded as unreconciled, and tools/test_ordered_gdn_recurrent.py keeps both measurements — a passing test and an expected failure that removes itself if the discrepancy is ours docs/m1-decisions.md (D8), the transcription in tools/ordered_gdn_recurrent.py, and tools/test_ordered_gdn_recurrent.py. 125 Python tests (2 expected failures)
2026-09-15 M0c: what 4-bit costs, measured (DC-031 under way). The install is 2.51× smaller than the checkpoint (1,813,618,928 against 4,548,285,948 bytes), quantizing 320 tensors to 7.71 bits/weight including scales and zero points and keeping 170 — the embedding, the head, the norms, the Gated DeltaNet's decay and convolution — at bf16 or fp32, because I3 says what decides a discrete outcome stays above four bits. The policy is data and a role missing from it stops the install rather than defaulting. On the frozen prompts, running the same contract against the install through the dequantizer: 28 of 29 next-token decisions survive, with one flip on the prompt whose margins are smallest, and the worst relative logit divergence rises from 2e-06 (fp32 engine against fp32 reference) to 1.6e-01 – 2.7e-01. That is the I3 hazard measured rather than described: one decision in twenty-nine changes while every numeric check still looks reasonable. Two findings from building it: the install had to carry the kept tensors as well or it is not a runnable model, and the four-bit codes are signed — reading 0b1000 as 8 instead of -8 shifts a whole group by sixteen steps and still produces plausible weights (the test caught it immediately) tools/quantize.py, tools/quant_policy.json, tools/measure_quantization.py, and the numbers in docs/m0c-quantization.md, which also states what M0c does not yet show — the engine does not read the install, and long-context behaviour is untested
2026-09-15 M1's validation model is pinned and its importer is written (DC-030 closed). Qwen/Qwen3.6-35B-A3B (Apache-2.0) is qwen3_5_moe: 40 layers (30 Gated DeltaNet + 10 full attention), 256 routed experts with top-8 and a shared expert, 26 shards and 67 GiB of bf16. The inventory came from the shard headers read by range request (tools/make_qwen36_fixture.py) — 1045 tensors, no weights downloaded — which is what an importer must be tested against. Counting from those shapes: 34.66 B total, 3.45 B active per token, so the brief's "35B / A3B" is right, and the reason it is 3.45 rather than 3.0 is that 1.02 B of the always-active part is the embedding and the untied head. A slip of mine is recorded with it: subtracting the experts of only the thirty Gated DeltaNet layers gave 11.26 B, because the experts are in all forty. The layout that matters: the routed experts are stacked (experts.gate_up_proj [256, 1024, 2048], gate and up fused), so the IR now carries source-layout roles (expert.stack_gate_up, expert.stack_down, expert.shared.scalar) alongside the per-expert roles L3's repack pass produces — an importer maps names and does not reshape, so the two are two roles rather than one role and a convention, and a stack holding only the gate is refused by the shape contract sources/DatacenterIR/Qwen3_5MoEImporter.swift, six tests, the 126 KB fixture, and docs/reference-qwen36-35b-a3b.md, which also lists the four things not to guess — the router's arithmetic first, because I3 makes its top-8 a discrete decision
2026-09-15 The reference is now measurable, and it is reproducible. tools/trace_capture.py runs one forward pass with hooks at every layer boundary — fp32 compute, eager attention, deterministic algorithms, pinned threads and seed, all recorded in the trace — with a --tiny mode that needs no download, which is how the plumbing was settled before a 4.5 GB fetch. Measured on that model: two captures are bit-identical, and one thread versus four produced the same bytes (still pinned, and re-measured on the real checkpoint, because a fact about a 164k-parameter model is not a fact about 2 B). The same run quantifies what a tolerance would have cost: bf16 differs from fp32 in every element, with the absolute divergence growing through the stack (1.2e-4 at the embedding → 2.8e-2 at the final norm) while relative error reaches 13285% on near-zero activations — unmeasurable, and blind to a top-k flip. The differ also now refuses traces from different reference stacks, which is not theoretical: this machine carries an unpinned transformers 5.16.1 in its system interpreter beside the venv's pinned 5.17.0 docs/m0-decisions.md D3 and docs/trace-format.md; the tests are the evidence — test_two_captures_are_bit_identical, test_thread_count_did_not_change_this_model, test_bf16_diverges_from_fp32_everywhere_and_grows_with_depth, test_mixed_reference_stacks_are_refused
2026-09-15 The int4 path is checked from the Swift side, and M1's prompts are frozen (DC-034 in progress). The tiny qwen3_5_moe fixture now carries an install (392 KB in total) and a golden for it, and two new tests assert the engine against it bit for bit including the router's decisions — the only test that reads a rank-3 quantized tensor in Swift, which matters because the Python reader sized that payload's code block from shape[0] last round, and a reader that does the same still reconstructs plausible weights. A second test asserts an install says its own family, so the loader does not need the caller to know what it was handed. The memory of that bug is now a generation-time check too: the fixture generator compares its hand-written configuration against the spec the engine emits, and that check would have caught the numExpertsPerToken drift that hid for two rounds — the same key mistake was in the generator, in the contract and in a test helper, and only the engine's own spec disagreed. tools/m1_prompts.json is frozen with real token ids (5, 33, 40, 60 and 67 tokens): long exceeds one delta-rule chunk, repeat makes several positions ask for the same experts, and the rest are deterministic continuations, so the gate measures throughput on prompts that exercise the paths rather than on one convenient string. Its sha256 prefix is 73d2d8a056ac1b88; changing it changes the gate the install golden in make_tiny_qwen36_checkpoint.py, two tests in Qwen3_5MoEForwardTests.swift, and tools/m1_prompts.json. 55 engine tests. What remains for M1's gate is the measurement itself, which needs the 35 B checkpoint: the download is at 18 GB of 67 GB
2026-09-15 The gate's numeric definition was pinned by a measurement, not by a preference (DC-006, R13). Before any kernel was written, the obvious question was asked: can a controlled accumulation reproduce PyTorch's matmul? On a real layer's shapes, an explicitly ordered fp32 sum differed from torch in 3544 of 4096 outputs (mean 17 ULP) while both sat 3.148e-07 from fp64 — so the difference is summation order, not accuracy, and it belongs to torch's BLAS rather than to the model. M0 therefore has two references: trace_capture.py (torch) as the semantic oracle, and the new tools/ordered_reference.py as the numeric contract, in which every sum has a stated order (ascending index, one product per step, no reassociation, no FMA). Validated on the real Qwen3-0.6B: the ordered forward's logits argmax agrees with torch at every position, mean |Δ| 8.6e-6 at the final norm, while the last bits differ from the first matmul onward. Finding this before the engine exists turned an unachievable gate into a defined one docs/m0-decisions.md D3, docs/trace-format.md, and six tests in tools/test_ordered_reference.py — including the one that asserts the divergence still exists, so the reasoning cannot quietly expire
2026-09-15 DC-017 — the farm was inventoried over SSH: all four nodes identical (Mac mini M2, 8 GB, macOS 27.0, Xcode 27, Swift 6.4, Python 3.14.7, Docker running, QEMU present, 81–107 GB free SSD), node 4 being this checkout's machine; both network paths measured, and no Thunderbolt bridge found configured The Testbed page's Farm inventory, built from each node's own answer rather than from the brief
2026-09-15 The engine's forward pass is bit-identical to the contract, on a real model (DC-023). sources/DatacenterEngine now reads safetensors (memory-mapped, bf16/f16/f32), builds its spec through the IR importer, runs the qwen3 dense forward dispatching on role rather than tensor name, and writes a trace; datacenter-trace is the executable and tools/check_engine_contract.py is the one-command gate. On Qwen/Qwen3-0.6B at revision c1899de2…, 8 tokens: 87 of 87 tensors byte-identical, whole-trace digest 101195ec6be8839c… computed independently by each implementation, 7,680,000 bytes of data.bin identical, 5.7 s at -O, peak resident set 2.43 GB. The project's own differ agrees. This is the first time the engine's wiring — not just its ops — has been checked: a transposed weight or a norm on the wrong side would have left every op-level test green. Two real defects were found by running it rather than by reading it: the safetensors header's __metadata__ block is not a tensor, and dtype names are upper case in real checkpoints tools/check_engine_contract.py reruns the whole claim; docs/trace-format.md records the numbers and the two traps
2026-09-15 DC-007 and DC-025 — the golden-trace harness exists and is proven able to fail: tools/trace_format.py (64-byte-aligned container, a sha256 per tensor, a whole-trace digest, and a refusal if the trace was edited after capture), tools/trace_diff.py (first-divergence localisation with the element index and fp32 ULP distance, plus discrete decisions compared separately and exactly), tools/make_synthetic_trace.py (a model-free deterministic fixture carrying a per-layer router decision), and 37 tests that CI runs on every push docs/trace-format.md and the tests themselves: a one-ULP change is located to the element, a later change is not reported, a transposed tensor is a shape finding rather than a byte diff, a flipped top-k is caught with every float identical, and a trace edited after capture is refused
2026-09-15 This family's attention and Gated DeltaNet are qwen3_5's arithmetic, verified mechanically rather than by eye. The reference contract's last open item was "matches by name and shape is not matches by arithmetic". tools/compare_reference_modules.py parses both transformers modules, renames Qwen3_5Moe onto Qwen3_5, and compares every top-level entity's unparsed body line by line: 25 entities are identical, and they are exactly the ones that matter — Qwen3_5GatedDeltaNet, Qwen3_5Attention (including the per-head [query \| gate] split), Qwen3_5TextRotaryEmbedding, torch_chunk_gated_delta_rule, torch_recurrent_gated_delta_rule, l2norm, causal_conv1d_fn, apply_rotary_pos_emb, eager_attention_forward, repeat_kv. Seven differ, all reviewed: the decoder layer substitutes SparseMoeBlock for MLP and unpacks its tuple, the three wrapper classes carry the mixture's parameter names, two output dataclasses differ, and Qwen3_5RMSNorm differs by a decorator — @use_kernel_forward_from_hub('RMSNormZeroCentered'), present in one family and absent in the other, with an identical body. That decorator name is the reference's own word for the convention M0b found the hard way (a 5 % error at layer 0 from assuming weight rather than (1 + weight)) — two independent routes to the same fact. The consequence is the one that matters for M1: the M0b kernels are reusable without modification, so M1's kernel work is the mixture and the streaming around it, not a second attention implementation. The tool takes an --allow-differ list and exits non-zero on anything unreviewed, so it re-runs after a transformers upgrade instead of being a one-off observation tools/compare_reference_modules.py, the verified table in docs/reference-qwen36-35b-a3b.md, and the command in AGENTS.md
2026-09-15 DC-016 — Core ML tooling installed: coremltools 9.1.dev1 on CPython 3.14.7 in .venv, pinned in tools/requirements-coreml.txt. It was first installed on 3.13 with the stable 9.0, then moved to 3.14 when the operator set the project's Python standard; 3.14 has no stable coremltools wheel, so the pin is an exact pre-release The reference machine ran the full round trip on 3.14.7: MIL program → .mlpackage → load → predict [3.0, 4.0, 5.0, 6.0], on CPU_ONLY and on CPU_AND_NE
2026-09-15 The cache bug is found, fixed, and the cache is now verified on the real model. The cached path diverged from the uncached one from the second token; the margins said it was a defect rather than a D8 difference, because both paths were confident in different tokens — their logits differed by more than the top-2 gap. Isolation in the order the evidence arrived: GatedDeltaNet.decodeStep against the sequence path is 1.6e-06 relative (rounding, not the cause); cached attention against sequence attention was 6.1e-03, and that test is what found it; after the fix it is 0.000e+00, bit-identical. The bug: Ops.orderedMatmul takes its weight as [out, k] — the layout of a Linear.weight — and attentionStep built both weight matrices as [k, out]. Every shape was right, every number plausible, and the pairs being multiplied were the wrong ones; the sequence path passing the key head as stored and the value head transposed is what gave the layout away. On the real model the cached path now gives tokens and margins identical to the uncached (11751,11,264,3177; 1.6400, 0.0972, 1.2532, 2.6414), at 16.5 s/step against 52.2 s/step. One measurement also narrowed a claim I had already written down: for a prompt inside one 64-position chunk the replay is bit-identical, because the chunked rule is the recurrence there — D8's divergence needs a second chunk, and is 3.2e-07 relative over seventy positions. Two smaller bugs were fixed on the way: generateCached consumed the last prompt token twice, and --cached was left in the positional argument list sources/DatacenterEngine/Qwen3_5Forward.swift (attentionStep's weight layouts), ModelCache.swift, the five tests in ModelCacheTests.swift, and the corrected write-up in docs/m1-decisions.md. 62 engine tests, green under -Onone and -O; 125 Python tests
2026-09-15 The Python side had the same sharding bug as the Swift side, and both are now fixed. The engine's reader was opened only the first of 26 shards; so did ordered_qwen35_trace.SafetensorsSource and quantize.build_install — meaning the contract would have computed a partial model's forward and the install builder would have written an install missing five sixth of the layers and reported success. Both now use one shared tools/safetensors_source.py, so the Python side cannot quietly disagree with the Swift side about what a sharded checkpoint is, and a sharded checkpoint with no index is refused rather than half-read. Three new tests need no Swift, so they run wherever the venv does: the contract's forward on the sharded fixture is byte-identical to the single-file one (tensors and decisions), the refusal is asserted, and an install built from the sharded checkpoint has the same data.bin bytes as one built from the single file — the payload must not depend on how the checkpoint was split. One assertion of mine was wrong in a way worth keeping: tensors and skipped in the manifest overlap (skipped is a subset), so adding them counted 61 tensors for a 36-tensor model tools/safetensors_source.py, the three call sites, and tools/test_sharded_checkpoint.py (now 5 tests). 121 Python tests. Download at 54 GB, 18 of 26 shards
2026-09-15 The gate measures memory too, and the download turned out to be working rather than stalled. DC-032's gate is a budget ("resident memory stays inside the configured budget") and the brief gives about 4.5 GB of usable memory per node, so the gate now records each engine run's peak resident set size from the platform's /usr/bin/time -l — an upper bound on the process, and the report says so, because on macOS the figure includes clean file-backed pages. Verified on the fixture: 0.01 GB. Separately, the 35 B download looked stalled at 65 GB and 22 of 26 shards for two checks; it was not — four .incomplete blobs were being written, the largest at 3.3 GB and growing. The lesson is small and worth keeping: a size that has not moved is not evidence of a stall when the client writes to temporary files, and the honest check is the incomplete-blob list rather than the directory total the memory measurement in tools/run_m1_gate.py and the table in docs/m1-gate.md. 121 Python tests
2026-09-15 M1's model runs end to end: the engine reproduces the contract on the tiny qwen3_5_moe checkpoint bit for bit, router decisions included. A checkpoint of 236 KB carrying the family's real naming (model.language_model., an untied lm_head), the real configuration nesting, two layers (one Gated DeltaNet, one full attention), eight experts with a top-2, a shared expert — and two key heads to four value heads, the asymmetry that the 2 B model could never exercise, so the grouped-query path runs rather than being bypassed. tools/make_tiny_qwen36_checkpoint.py writes it and its golden, and it checks its own golden against the reference module before writing, so the fixture cannot be golden for the wrong model (measured: max |Δ| 1.4e-06 against scale 3.06). Five tests then assert the engine against it: every captured tensor's bits, the greedy continuation, and — as its own assertion — the router's decisions at every layer, plus that no discrete decision is carried as a tensor, because a tolerance cannot express "the same experts". One implementation now serves both qwen3_5 families: the reference branches inside its decoder layer between a dense feed-forward and a mixture, compare_reference_modules.py proved the rest of the arithmetic identical, and a second copy would be a second thing to keep in step. The install path now dispatches on the spec's family, so an artifact says which forward pass reads it; a qwen3 install is refused rather than guessed at. The one thing this round did not solve is the reason M1 exists: the forward materialises a whole layer's experts, which is 3.2 GB in fp32, so it runs on the tiny fixture and cannot yet run on the real model — that is DC-032, now the next task tools/make_tiny_qwen36_checkpoint.py, tests/DatacenterEngineTests/Qwen3_5MoEForwardTests.swift, the ForwardResult in ForwardPass.swift, and the mixture branch in Qwen3_5Forward.swift. 46 engine tests, green under -Onone and -O. Also fixed: the fixture generator silently used a stale release binary for the spec, the same trap that has bitten end-to-end runs before
2026-09-15 The qwen3_5 tower runs in Swift, bit-identical to the contract, with the weights loaded one layer at a time (DC-029). 2 B parameters are 8 GB in fp32 and the nodes have about 4.5 GB usable, so nothing is loaded whole: the embedding is read a row at a time (a token needs one row of a [248320, 2048] matrix), the tied head is computed in blocks of vocabulary rows, and each decoder layer's tensors are read, used and released before the next starts — peak residency is one layer, ~330 MB. Qwen3_5ForwardTests checks it bit for bit against the contract on a committed 157 KB checkpoint built from this family's real geometry and naming: three Gated DeltaNet layers and one full-attention layer, a partial RoPE, a tied head, the real nested config file. Every captured tensor matches exactly and the logits' argmax matches as a discrete decision. Two findings from building it: the v5 config keeps rope_theta and partial_rotary_factor inside rope_parameters, so reading a top-level key left the base at 10 000 instead of 1e7 and the rotation factor unset; and Qwen3_5TextConfig has no attn_output_gate attribute at all while the attention module hardcodes the doubled query and never reads a flag — the flag describes the checkpoint, it does not switch anything sources/DatacenterEngine/Qwen3_5Forward.swift, tests/DatacenterEngineTests/Qwen3_5ForwardTests.swift, and the fixture built by tools/make_tiny_qwen35_checkpoint.py. Left for the next round: the same run against the real 2 B checkpoint, and a CLI that dispatches on the checkpoint's own model_type
2026-09-15 DC-082 — the documentation sync rule was recorded and applies from this date This page, "How to read this page"
2026-09-15 CI caught the M1 gate test on a runner that cannot build this package (DC-036, again). The Markdown job on GitHub runs the whole tools/ suite, and the two new tests build the engine — which a swift-tools-version:6.4 package cannot do on the macos-26 image, whose toolchain is 6.3.3. The tests now read swift --version and skip below 6.4, the same way the Swift CI job prints its toolchain and gates only at 6.4 or above, so the suite is honest on every runner instead of green on one. Worth recording that this is the second task the toolchain gap has cost (DC-036 is still open), and that the failure came from a test I had verified locally — a local pass says nothing about the runner's toolchain the version gate in tools/test_run_m1_gate.py; the local toolchain is 6.4 and the tests run
2026-09-15 The engine generates coherent text, and every token matches the contract (DC-023 closed). datacenter-generate runs greedy decoding — the whole sequence re-run each step, because M0 deliberately has no KV cache and a cache would be a second numeric path before the first one is trusted — and records the generated tokens as a discrete decision in the trace. On Qwen3-0.6B with the prompt "The capital of France is": the engine produced Paris. The capital of Italy is Rome, 8 of 8 token ids identical to the contract, and the trace identical (87 tensors, 1 discrete decision, matching digests). This is I3 exercised for real for the first time: the sampled token is the most consequential discrete decision in the model, a 1-ULP logit difference can flip it, and it is compared as an index set rather than as a number. Also fixed en route: TraceWriter.Discrete had no public initialiser, so it could not be constructed from another module at all — a library type that only its own module can use is a type with no callers tools/check_engine_generation.py reruns the claim and prints the decoded text; five new tests pin the sampler's window and tie-break, including that two traces differing only in generated tokens have different digests
2026-09-15 The whole qwen3_5 text tower now matches the reference model, and doing so found a trap a shape check could not (DC-029). Qwen3_5RMSNorm:841 is weight-offset: its parameter is initialised to zeros and the multiply is by (1 + weight), so a zero weight is the identity and a checkpoint's stored values are offsets from one. The same file contains the ordinary kind inside the Gated DeltaNet, and the qwen3 family uses the ordinary kind too — three norms, two conventions. Assuming they agreed gave a 5 % error at layer 0 while every shape stayed right, and it was caught by comparing against the reference module rather than by reading the code. Two independent confirmations are recorded: a fresh module with an all-zero weight returns a non-zero output, and the real checkpoint's layers.0.input_layernorm.weight is centred near zero (mean 0.0956, min −0.2051, max 1.0859) where a plain weight would be centred near one. tools/ordered_qwen35.py now carries the offset convention, and 9 tests check the whole tower against transformers' own model: the final hidden state at two seeds, a 70-position sequence spanning a chunk boundary, every layer individually so error cancellation cannot hide a wrong one, the tied head's logits and argmax, the partial RoPE's width and pass-through, and the offset itself tools/ordered_qwen35.py, tools/test_ordered_qwen35_model.py, and the convention table in docs/reference-qwen35-2b.md; next is the same forward in Swift, which needs per-layer weight loading because 2 B parameters do not fit in fp32 on an 8 GB node
2026-09-15 The Swift gate exists, and its first run found the toolchain gap (DC-036). .github/workflows/swift.yml failed with the exact diagnosis it was written to produce: package 'tinytitan_datacenter' is using Swift tools version 6.4.0 but the installed version is 6.3.3. A probe step then listed the image's Xcodes — 26.0, 26.0.1, 26.1, 26.1.1, 26.2, 26.2.0, 26.3, 26.3.0, 26.4, 26.4.1, 26.5, 26.5.0 and no 27 — so GitHub-hosted runners cannot build this project at all today. Options weighed: lower the manifest to 6.3 (contradicts the project standard and silently permits an older toolchain), move the gate to a self-hosted runner on the farm (the right long-term answer, needs a registration token and changes the operator's infrastructure), or make the job self-documenting. It now prints the toolchain, gates when it is 6.4+, and otherwise warns by name; it switches itself on when an image ships Xcode 27. Until then the real gate is swift build && swift test on the farm The workflow and its runs; DC-036 carries the gap
2026-09-15 M1's correctness claim is measured on the real 35 B checkpoint: the engine reproduces the contract byte for byte, router decisions included. capital (5 tokens) through the whole stack — sharded reader, importer, streaming expert provider, mixture, head — gives 83 tensors, 40 discrete decisions and digest b8c976c5e7ba8816… on both sides, and trace_diff reports IDENTICAL with the discrete entries checked as their own kind. The engine took 45.6 s against the contract's 461.8 s — 10.1× — which is the first honest fp32 figure on this hardware. Three numbers that matter more than the headline: peak resident memory 4.31 GB against the brief's ~4.5 GB of usable memory per node, which is at the budget edge and is a measured argument for the brief's own F_NOCACHE/O_DIRECT rule for expert slabs (not yet implemented); cache hit rate 0.0000 across 2240 expert requests in 40 layers, because a layer's slot bank is built as the layer loads and dropped with it, so nothing survives a token; and the traffic itself — 40 layers × 28 distinct experts × 2 projections, consistent with 5 positions × top-8 = 40 pairs, ≈ 14.1 GB of decoded fp32 or ≈ 7.05 GB as the bf16 the shards store. One piece of my own arithmetic was wrong and worth recording: I checked the counters expecting top-2 and got a number four times too large — the real model is top-8, as its config says, and the fixture I had been reasoning from is not docs/m1-gate.md (first result), .build/m1-capital-{engine,contract}. Two small fixes follow: the slot cache's rowsRead counts elements while the counting provider counts rows, so the unit is mislabelled and the gate should report bytes; and the full five-prompt gate run plus generation throughput is the next measurement
2026-09-15 Where the 52 s/token goes, and why a cache is a second numeric path rather than an optimisation. Measured on the real model: one position costs 19.1 s (with ~2 GB read from the SSD) and the five-position prefill 45.6 s, so the marginal cost is ~6.6 s per position — about 70 % compute and 30 % I/O — and the generation's mean step of 6.5 positions is 52 s, which is that arithmetic exactly. So a cache turns a decode step from ~6.5 positions into one and wins ~8× at this length, growing with context; and the compute is the larger term, which is what makes the kernels (DC-033) the biggest single lever. Then the design question, answered from the reference rather than from taste: Qwen3_5MoeGatedDeltaNet.forward:625 switches functions — torch_recurrent_gated_delta_rule when use_precomputed_states and seq_len == 1, torch_chunk_gated_delta_rule otherwise — so the reference itself does not produce the same bytes with and without a cache. The chunked rule groups its sums over 64 positions and the recurrent one accumulates a step at a time. Three decisions follow: M1's bit-identity claim belongs to the uncached path (which is what the gate measured); the cached path is validated against the oracle's own cached path with the discrete decisions still asserted exactly, because I3 does not relax when the path changes; and I1 is not weakened — the same prompt in the same mode is still deterministic, and what I1 never promised is that two different numeric paths agree the transcription with line numbers in docs/reference-qwen36-35b-a3b.md, the measurements in this row, and .build/m1-one-engine/metrics.json (640 expert requests for one position: 40 layers × 8 experts × 2 projections)
2026-09-15 The engine reads sharded checkpoints, without which M1 cannot run at all. The 35 B model is 26 shards with an index, and the loader opened only the first shard — producing a checkpoint that looks complete and is missing five sixth of its layers, which is precisely the failure a shape check cannot see. ShardedSafetensors reads the index, opens shards lazily and keeps the mappings (an mmap costs no resident memory until a page is touched, and re-opening per read would be a syscall storm on the path where a token fetches eight experts from eight places). Both loaders now go through one SnapshotWeights.open, so a sharded checkpoint and a single file present the same inventory to the importer and the file format stays out of L2. Two tests: splitting the tiny fixture into three shards and requiring the engine's trace to be identical to the single-file one, and requiring a sharded checkpoint with no index to be refused rather than half-read. One test bug worth recording: I named a helper run, which shadowed TestCase.run and broke the runner; and the refusal test initially passed against a stale release binary because tests run in name order and only one of them built — the build now happens once in setUpClass sources/DatacenterEngine/ShardedSafetensors.swift, the factory in SnapshotWeights.open, and tools/test_sharded_checkpoint.py (2 tests). 118 Python tests. The download is at 48 GB with 15 of 26 shards
2026-09-15 M1's gate exists as one command and is verified on the fixture rather than waiting for the 67 GB model. tools/run_m1_gate.py runs all three of the gate's parts on one checkpoint: correctness (every frozen prompt, engine trace against contract trace through trace_diff, which compares the discrete decisions exactly as well), throughput (greedy generation with its own timing parsed into the report), and the cache hit rate (the engine's expert counters, now written beside each trace in metrics.json — beside, not into the manifest, because the manifest carries the digest and a counter that varies between runs would make I1's comparison impossible). Two new pieces were needed to get there: tools/ordered_qwen36_trace.py, the mixture's contract CLI, whose traces come out byte-identical to the engine's with equal digests and the decisions checked, and a family guard, so M1's gate refuses a qwen3_5 checkpoint rather than measuring something else and labelling it M1. tools/test_run_m1_gate.py drives the whole script against the tiny fixture, so the gate is not untested code waiting for the one input that matters. One finding already, from the tiny fixture: the cache hit rate is 0.0000, because the slot bank belongs to one layer and is built as that layer loads, so a hit needs two positions in the same layer of the same pass to choose the same expert — and nothing is reused across tokens. Whether the real model's eight-of-256 routing repeats enough to earn the bank's memory is what the gate exists to answer, and a cache that survives a token would be a design change to record as one. A bug in the instrument itself: if seconds is false for a legitimate 0.0, which silently turned a measured run into None tools/run_m1_gate.py, tools/ordered_qwen36_trace.py, tools/test_run_m1_gate.py (2 tests) and docs/m1-gate.md. 116 Python tests. The measurement itself still needs the checkpoint: 25 GB of 67 GB fetched
2026-09-15 The qwen3_5 importer, checked against all 632 tensors of the real checkpoint (DC-022 closed). Qwen3.5-2B's text tower is model.language_model.* — and only 320 of its 632 tensors are text. The mapping was written from the safetensors header rather than from a reading of the reference, which settled four things the documentation had left open: full attention is exactly index % 4 == 3 (3, 7, 11, 15, 19, 23); q_proj is [4096, 2048] because attn_output_gate puts the gate in the same tensor as [query \| gate]; A_log and the Gated DeltaNet's norm are stored fp32 among bf16 neighbours; and there is no lm_head at all, the embeddings being tied. Two IR changes fell out and both are load-bearing: attn.q's shape contract is now gate-aware, and 312 non-text tensors (297 vision, 15 MTP) are excluded by a named prefix with a reason rather than skipped — the tests assert every tensor is either mapped or excluded, because a third category is where a silently unused weight would hide tests/DatacenterIRTests/Qwen3_5ImporterTests.swift (8 tests) and the 47 KB inventory fixture; docs/reference-qwen35-2b.md records what the checkpoint settled and, still explicitly, the six things it did not
2026-09-15 The pinned 2 B model runs on one node, and its decisions match the reference implementation (DC-029). The engine streamed Qwen/Qwen3.5-2B at revision 15852e8c… — 8 tokens in 19.3 s, peak resident set 3.41 GB, most of it clean file-backed pages of the memory-mapped checkpoint rather than working set, because one decoder layer is resident at a time. Against transformers (the semantic oracle, a different summation order, so byte-equality is impossible by D3/R13) the numbers differ by at most 4.1e-6 relative after 24 layers — the shape of a rounding-order difference and nothing else — while all 8 discrete decisions match exactly. The smallest top-1 margin on that prompt is 0.2192, three orders of magnitude above the divergence: that is why the decisions survived, and it is also why they are asserted separately (I3) rather than inferred from the numbers being close. tools/compare_engine_to_oracle.py reports the two claims apart, so a future run cannot pass by being numerically close and discretely wrong. Also added: ModelLoader dispatches on the checkpoint's own model_type, so one binary runs both families and the operator does not choose. What remains for M0b's bit-exact half is the contract running the 2 B model — which needs the same per-layer streaming on the Python side, since it currently loads 8 GB of fp32 The commands are in AGENTS.md and the numbers in docs/trace-format.md; tools/compare_engine_to_oracle.py reruns the comparison
2026-09-15 Three defects found by running the tool, not by reading it. (1) A forward pre-hook that returns non-None replaces the module's arguments, and the weight-loading hook was returning load_into's tensor count — the first real capture failed with an int where the hidden state should have been. (2) Assigning a meta tensor to param.data is refused by PyTorch's variable hooks, so releasing a layer replaces its parameter table instead of mutating it. (3) The disk path named its determinism record record, the same name as its tensor-recording closure, so the manifest tried to serialize a function. All three are fixed, and the first two are now regression-tested (test_disk_capture_matches_the_resident_capture, test_disk_capture_refuses_a_tensor_no_module_claims) 48 tests, green under both the pinned venv and the system interpreter
2026-09-15 The IR is defined, implemented and tested against a real checkpoint (DC-010, DC-012). docs/ir-schema.md is the contract: the spec file's shape, a closed role vocabulary with a shape contract per role, per-role quantization and sharding tables, provenance, and the seven diagnostics a spec can raise. docs/repository-layout.md records the layout. In Swift, sources/DatacenterIR/ holds the types, the validator and the qwen3 importer, with 13 tests. The importer was tested against the real inventory of Qwen3-0.6B — 311 tensors with their shapes, dumped from the safetensors header and committed as a fixture — and maps every one of them, with QK-norm landing on [headDim] and not [hiddenSize] The tests are the evidence. A mapping tested only against names we invented would be a mapping of our own assumptions, and this one was written to fail on the real file instead
2026-09-15 DC-038 closed: the Swift Gated DeltaNet expands the key and query heads, and the fix is checked bit-for-bit. The golden vectors gained an asymmetric case — two key heads to four value heads, where the 2 B model's sixteen and sixteen could never tell whether the grouped-query repeat was there — and GatedDeltaNetTests now asserts it against the contract bit for bit. The vector pins the order of the expansion too: consecutive ([k0, k0, k1, k1]) as repeat_interleave produces, and an interleaved reading pairs different heads and gives different bits. The kernel also asserts that the value head count is a multiple of the key head count, which the reference does not check — there a non-multiple would silently repeat the wrong number of times. A second test I wrote for the ordering was deleted rather than repaired: it asserted shapes with arithmetic I had got wrong and could say no more than the bit-exact vector already does, and a test that adds nothing is worse than no test because it has to be maintained sources/DatacenterEngine/GatedDeltaNet.swift, the gdn_asymmetric group in contract-vectors.json, and tests/DatacenterEngineTests/GatedDeltaNetTests.swift. 41 engine tests, green under -Onone and -O
2026-09-15 M1's mixture of experts is transcribed and checked against the reference (DC-037). Read from Qwen3_5MoeTopKRouter:884, Qwen3_5MoeExperts:845 and Qwen3_5MoeSparseMoeBlock:903, and now recorded with line numbers in docs/reference-qwen36-35b-a3b.md. Four things the reading settled, each of which a guess would have got wrong or left open: the softmax is fp32 while everything around it is bf16; the top-k is taken on the probabilities, not the logits; the chosen weights are renormalised unconditionally (the implementation never consults norm_topk_prob); and the fused stack is [gate \| up] with the gate first, which the shape could not say. Two things are decisions rather than transcriptions and are recorded as such: ties break to the lowest expert index, because torch.topk promises nothing about equal probabilities and I3 needs a comparable index set; and the routed accumulation runs in ascending expert index, which the reference does (it iterates nonzero, which is sorted) and which is therefore also what D4's ring reduction will do — the single-node contract and the distributed one agree by construction rather than coincidence. tools/ordered_moe.py implements it and seven tests check it against the reference's own module, the top-k index set separately from the numbers. One trap found by testing: Qwen3_5MoeExperts allocates with torch.empty, so a from-scratch tiny MoE block holds uninitialised parameters — both sides computed NaN and the first run "passed" on two tests for the wrong reason, which is why the tests now assert finiteness before comparing tools/ordered_moe.py, tools/test_ordered_moe.py, and the transcription in docs/reference-qwen36-35b-a3b.md; what remains before kernels is the line-by-line comparison of this family's attention and Gated DeltaNet with qwen3_5's, which the doc still lists as unverified
2026-09-15 DC-027 — the reference contract was written, first for qwen3 dense and then, after D1 moved the milestone to Qwen3.5-2B, for that family too: docs/reference-qwen3-dense.md (17 steps of a dense layer, all dtype boundaries) and docs/reference-qwen35-2b.md (the Gated DeltaNet wiring), both read from transformers v5.17.0 with file sha256 recorded at revision 70d244cc and 15852e8c respectively. Between them they name the traps the brief warns about rather than leaving them to be guessed: the per-head QK-norm applied before RoPE, the norm casting back to bf16 before the weight multiply, softmax as the only fp32 island, rotate-half RoPE with bf16 cos/sin, the tied head, the legacy rope_theta key the v5 reference translates but the checkpoint still ships, and the gated norm that activates its gate in fp32 The documents themselves, cited by line number; torch==2.14.0 and transformers==5.17.0 pinned in tools/requirements-reference.txt so a trace stays comparable to the build that produced it
2026-09-15 The reference's own two Gated DeltaNet paths disagree with each other by 3.4 relative, and the cache work stops here until that is settled. Building the decode path test-first (tools/test_ordered_gdn_recurrent.py) produced a transcription whose recurrent rule agrees with our chunked rule to 3.2e-07 relative — and our chunked rule is the one M0 validated against the reference's chunked function. But the reference's own recurrent function, on identical inputs in the pinned version, differs from the reference's own chunked function by 3.4 relative: a complete divergence rather than rounding. Two lines were genuinely missing from my first attempt and are now in — the axis convention ([batch, length, heads, dim], which broadcasts happily and returns wrong numbers when transposed) and query = query / (query.shape[-1] ** 0.5), applied always after the l2norm — and with those, our paths agree with each other. What is left is not ours to fix by guessing: the working agreement says "when a reference implementation is ambiguous, say so and ask", and this is that case. The two reference-comparison tests are expected failures carrying the reason, so the suite stays meaningful and the marker removes itself if the discrepancy is ours. Consequence for M1's cache: the decode arithmetic has two candidate meanings, so the cache cannot be called faithful to either by assumption. What is true either way is that the paths are algebraically equivalent — a cache is a different numeric path, not a different model tools/ordered_gdn_recurrent.py (the transcription), tools/test_ordered_gdn_recurrent.py (4 tests: 2 pass, 2 expected failures), and the measurement written up in docs/reference-qwen36-35b-a3b.md. 123 Python tests
2026-09-15 The Gated DeltaNet kernel runs in Swift, bit-identical to the contract (DC-029). sources/DatacenterEngine/GatedDeltaNet.swift implements the whole layer — the projection, the depthwise causal conv, the per-head gates, the chunked delta rule with its padding and sequential chunk scan, the gated norm and the output projection — and GatedDeltaNetTests asserts golden bit patterns from tools/make_contract_vectors.py for a single-chunk case and a 70-position, two-chunk case, because the scan over chunks only runs in the second. It passes under -Onone and -O: a compiler that contracted a * b + c into an FMA would break every invariant here silently, so "optimisation does not disturb the numerics" is a test rather than an assumption. Two real defects came out of the cross-language check: the gate and decay arrays were written once per position instead of once per head, and the contract itself carried a second silu with a different branch structure, which appeared as a 1-ULP disagreement and is now imported from one place sources/DatacenterEngine/GatedDeltaNet.swift, tests/DatacenterEngineTests/GatedDeltaNetTests.swift, and the vectors in tests/DatacenterEngineTests/Fixtures/contract-vectors.json. What remains for M0b is the family's whole-model forward in Swift plus per-layer weight loading, since 2 B parameters do not fit in fp32 on an 8 GB node
2026-09-15 Correction: R11's bf16 example was wrong. It quoted two logits at 5.424064 and 5.413122 collapsing to the same bf16 value; those numbers came from mantissa truncation, which is not what a bf16 cast does. Re-measured with round-to-nearest-even over 20,000 random 256-way routers: 922 (4.61%) flip the top-8 set, e.g. 5.363673687 and 5.385571480 both become 5.375. The conclusion is unchanged — D5's fp32 routers stand — but the evidence is now accurate, and the pair is an executable fixture rather than a sentence The corrected R11 row, docs/m0-decisions.md D5, and test_bf16_rounding_reproduces_the_router_hazard. Found by writing the test for a claim that had only been asserted
2026-09-15 M0a is complete: the reference is captured two ways and the ways agree. trace_capture.py gained the layer-by-layer path (--from-disk): the model is built on the meta device and one decoder layer is loaded, used and released at a time. On the real Qwen3-0.6B the resident and streaming paths are bit-identical (87 tensors, same digest). On the real M0 model Qwen3.5-2B — 18 Gated DeltaNet layers, ~9.2 GB of fp32 weights — two independent runs are bit-identical (75 tensors, digest 72793ffc…) at a peak RSS of 2.78–3.17 GiB; the resident path could not have produced that trace on this hardware at all. The trace records delta_rule_path: chunked, 18 linear-attention layers and optional_kernels: {causal_conv1d: false, fla: false}, so a consumer knows the reference used its own PyTorch fallback rather than a fused kernel docs/trace-format.md, docs/m0-decisions.md D3/D6, and the runs themselves. The tool also refuses two things rather than guessing: a checkpoint tensor no module claims, and a tie_word_embeddings: true config whose shipped lm_head.weight differs from embed_tokens — which Qwen3-0.6B ships separately, so the check was not hypothetical
2026-09-15 The mixture of experts runs in Swift, bit-identical to the contract. sources/DatacenterEngine/MixtureOfExperts.swift implements the router (fp32 softmax, top-k on the probabilities, unconditional renormalisation), the stacked experts (gate first in the fused stack, accumulated in ascending expert index) and the shared expert with its scalar gate — and MixtureOfExpertsTests asserts golden bit patterns from tools/ordered_moe.py for the block's output and the renormalised weights, with the chosen experts asserted separately, in order because I3 says so and that is the assertion quantisation breaks first. It passes under -Onone and -O, which is the check that the compiler is not contracting a multiply-add into the fused operation the contract forbids. Two tie cases are tested directly rather than through a vector that happens to contain no ties: Swift's sorted(by:) is not a stable sort, so the comparator carries the index and the tie-break is a total order rather than a hope. Seven new tests, 40 engine tests in total sources/DatacenterEngine/MixtureOfExperts.swift, tests/DatacenterEngineTests/MixtureOfExpertsTests.swift, the vectors in contract-vectors.json, and the kernel section of docs/reference-qwen36-35b-a3b.md. What remains for M1's forward pass is assembling the family: the mixture behind a per-layer streamed loader, the decoder layer, and the router's indices recorded as a discrete decision in the trace
2026-09-15 The engine's ops reproduce the contract bit-for-bit, in a different language (DC-023, first half). Before writing a forward pass, the question that decides whether cross-language bit-exactness is possible at all was measured: numpy's float32 sin/cos differ from Swift's in 12–18% of cases, and Swift's differ from libm's own sinf/cosf in ~1% — so exp (0/3510) and pow (0/64) were safe while RoPE's trigonometry was not. The fix is a contract sentence rather than a tolerance: transcendentals are computed in double precision and rounded to Float, which agreed in 0 of 6000 samples. sources/DatacenterEngine now implements ordered matmul, ordered sum, RMSNorm, silu, sigmoid, softmax, the RoPE tables and rotate-half, and tools/make_contract_vectors.py emits golden bit patterns that nine Swift tests assert. All 22 Swift tests pass. Audited by mutation, because a gate that cannot fail is not a gate: reversing the accumulation order fails the matmul test, and swapping in native cosf/sinf fails the RoPE test, while exp is provably indifferent (0/3510) — which is exactly why the rule is stated uniformly instead of as a list of exceptions The fixture and the tests are committed; the measurements are in D3 in docs/m0-decisions.md
2026-09-15 The D1 model verified before anything is built on it — Qwen3.5-2B is not a conventional transformer (24 layers, 18 Gated DeltaNet + 6 full attention, depthwise conv k=4, chunk_size=64, beta=sigmoid(b), g=-exp(A_log)·softplus(a+dt_bias), gated RMSNorm), it is a vision-language checkpoint whose text tower is the part we need, and it is Apache-2.0. The conflict with M0's purpose is recorded as R12 and answered with the M0a/M0b/M0c staging docs/reference-qwen35-2b.md, read from transformers v5.17.0 (modeling_qwen3_5.py, sha256 recorded) at checkpoint revision 15852e8c…
2026-09-15 The decode cache is built, measured at 3.2× on the real model — and withheld from the gate because the real model says it is wrong. ModelCache.swift decodes one position per token against a per-layer state: GatedDeltaNet.decodeStep for the recurrence (window plus state) and attentionStep for the cached attention, with datacenter-generate --cached to measure. Cached: 66.1 s for 4 steps (16.5 s/step) against the uncached 208.9 s (52.2 s/step) at a five-token prompt — 3.2×, and it grows with context. On the tiny fixture everything holds: cached and uncached produce identical tokens, the router's decisions agree 2 of 2 layers exactly, and the replay lands on the prompt's last position within 2.2e-03 relative. But on the real model the two paths agree on the first token and diverge from the second (11751, 13, 561, 6511 against 11751, 11, 264, 3177), while the reference's own cached and uncached greedy generation produce identical tokens on the same fixture. That asymmetry is what makes it a suspected bug rather than a D8 numeric-path difference, and it is why the speedup is not being banked: M1's gate continues to rest on the uncached path, which is bit-identical to the contract. The isolating experiment is named — per-layer hidden_out, cached against uncached, so the first layer that disagrees names the bug. Two bugs were already found and fixed on the way: generateCached consumed the last prompt token twice (the replay already had), and --cached was left in the positional argument list so every cached invocation printed the usage line sources/DatacenterEngine/ModelCache.swift, attentionStep in Qwen3_5Forward.swift, the four tests in ModelCacheTests.swift, and datacenter-generate --cached. 61 engine tests, green under -Onone and -O
TinyTitan Datacenter

Project

Design

Development

Sister project: TinyTitan

Clone this wiki locally