Skip to content

Project Tracker

Pummelchen edited this page Sep 18, 2026 · 259 revisions

Current status

Phase P2 — M1 (P0 in progress; P1's M0 gate passed)
M0 Gate passed — Qwen/Qwen3.5-2B revision 15852e8c…, three frozen prompts, 40,683,520 bytes of trace data identical to the contract and every discrete decision matching the reference (docs/m0-gate.md)
M1 Gate passed — and now verified against a contract reading the same weights (2026-09-17). trace_diff between the engine and a contract run on the install itself reports IDENTICAL — 83 tensors, 0 differing elements, 40 discrete decisions, matching digests b0d382dbabf36df0…. The claim is restated in tools/milestones.json to that falsifiable form, because the earlier comparison had two inputs: the engine read an install and the contract read a bf16 checkpoint, and D55 showed the whole divergence was that difference — the DeltaNet's three projections and attn.q/k/v/o are int4 in the install, which moves layer 0's attention output by 0.0156 max / 24.5% median relative. Against the checkpoint contract the pair differs by 40 discrete decisions and 1 float tensor, declared and tracked in DC-112. Engine measurements that stand: 0.108 tok/s cached and 0.0374 uncached generation with identical tokens, a hit rate of 0.0000 over 2,240 requests (structural), and 348.6 MB peak memory
M2 Gate passed, and re-established on 2026-09-17 against the current binaries. A 256-expert plan over two contiguous halves, one node on a peer machine over TCP and one here, each reading its own install half (2.05 GB apiece), 40 reductions per node over 797 and 803 terms: trace_diff reports IDENTICAL — 83 tensors, 0 differing elements, 40 discrete decisions, matching digests b0d382dbabf36df0… for both nodes against the single-node reference. This matters because the engine has changed since the original run — the GPU unpack became the default and every contract matmul went through a chooser — and both are asserted bit-identical, which is exactly the sort of assertion that wants re-checking rather than trusting (D66).
Open tasks 36 across the phases below; closed work and the dated record are on News
Language standard Swift 6.4 on Xcode 27 with swift-tools-version:6.4 — the only supported toolchain, enforced by tools/check_toolchain.py; Python 3.14
Tests 266 Swift, 427 Python, both run on every push — and nothing is skipped: the Metal kernel tests run on the node's GPU since D34
Farm Four identical Mac mini M2, 8 GB each, macOS 27.0, Xcode 27, Python 3.14.7
Development link 1 Gbit Ethernet: 0.49–0.64 ms round trip measured; the mesh VPN path measures 1.4–1.8 ms

Phases at a glance

Phase Milestone Objective Gate Status
P0 Foundation Decisions, harnesses and conventions ADRs, harness spec and conventions written In progress
P1 M0 Staged: harness (M0a), Qwen3.5-2B bf16 (M0b), 4-bit (M0c) Bit-matches reference golden traces at each stage Done
P2 M1 Qwen3.6-35B-A3B, single node, 4-bit, SSD-streamed Correct output + recorded tok/s baseline Done
P3 M2 2 nodes, expert-parallel Bit-identical to M1 Done
P4 M3 4 nodes ≥3x the M1 tok/s In progress
P5 M4 DeepSeek-V4.1-Flash Matches reference at 128K context Planned
P6 M5 Qwen3.8-Flash-Next Same protocol as M4 Planned

How to read this page

  • Task IDs (DC-nnn) are stable and are never reused or renumbered. A cancelled task keeps its ID and is marked Dropped with the reason.
  • Status is one of Planned · Next · In progress · Blocked · Done · Dropped. Next means it is the immediate queue, not that it has started.
  • "Done when" is the evidence, not the activity. A task is done when the named artifact or measurement exists, and the row says where it is.
  • A gate is closed only by measurement. When a status here disagrees with a gate on the Roadmap, the measurement wins and this page is corrected.
  • Documentation sync rule (DC-082). A commit that changes behaviour, the plan, or a decision updates the README, the affected wiki page and this tracker in the same push. A tracker row that has drifted from the repository is a defect.

P0 — Foundation

Decisions that are expensive to reverse. Each ADR records the alternatives considered and why the others were rejected.

ID Task Depends on Done when Status
DC-127 Wire the fused int4 matmul into the expert path, so a resident slab never becomes fp32 D108 built and measured the kernel: it is bit-identical to the split and 1.23x / 1.45x / 2.36x faster on the real shapes, worth about 6.6% of a step by itself. The reason to wire it is not that 6.6% — it is that it is the precondition for residency: a resident packed slab can only avoid the fp32 materialisation if the matmul consumes the packed form. The obstacle is the interface, not the arithmetic: WeightSource exposes only fp32 (tensor, rows, rowsStreaming), so the provider must be able to hand over a packed payload and its entry, and ExpertSlotCache/ExpertBank — which hold [Float] slices and whose whole point is to be hit — have to be reconciled with a path that has no fp32 to cache A decode step whose expert matmuls run on the fused kernel with the trace digest unchanged and the step measured lower alternated on one binary, gated by the M1 install trace; the wiring lands only if the gate passes. Partly landed and measured (D109): packedRows, gateUpProduct/downProduct and preloadPacked exist and are tested, and with SHARD_GPU_INT4_EXPERTS=1 plus a 1 GB slab cache the reads fall 8.46 → 5.73 GB and the step returns to 1.575 s — a wash against the 1.69 s control, because 640 dispatch-and-wait pairs per token cost what the kernel saves. The next step is one dispatch per layer, not one per expert. Landed and measured (D110): the kernel is 3.6-4x faster on wide loads and the path is the default at a 256 MiB slab cache — 1.465 s/step, 0.682 tok/s, lower peak RSS than the split path. Batching per layer is still the open part: 640 synchronous dispatches remain, and at 0.3-0.9 ms each that is 200-500 ms a step Open
DC-126 A cache of packed slab payloads, so a read can be avoided rather than hidden D105 concluded the device is saturated, so the lever is reading fewer bytes. InstallFile.SlabCache holds expert row ranges keyed by name and range as stored — four bits per weight rather than the thirty-two the decoder produces — and a hit skips the three preads altogether. Measured, and it loses on this node: at 512 MB it never hits at all (0.0%, 536 MB held) because a slab is evicted before the next token asks; at 768 MB it hits 36.1% and reads 0.73 GB/step; at 1024 MB it hits 45.0% and reads 0.65 GB/step against 1.08 — mix.read 0.80 → 0.54 s — and the step gets worse (1.931-2.022 against 1.735), because load nearly doubles (0.363 → 0.591-0.650) and head grows (0.261 → 0.348-0.411) on a machine already swapping (1.94 of 3.07 GB). The mechanism is validated and the node cannot afford it; the default is 0 and SHARD_SLAB_CACHE_MB is the knob. Digest 89d654ff54b0fd03 in every arm SlabCacheTests (5), the A/B in D106, SlabCacheTests.swift Done (negative)
DC-125 A deterministic regression test for the StreamHandle race DC-124 was found by correlation and closed on a failure rate (3 of 4 → 0 of 5), which is a measurement rather than a test: a race that only shows under timing pressure has no deterministic reproduction, and the fix deserves one that fails loudly if the lock is ever removed A test that calls a checkpoint's rowsStreaming from many threads at once — enough iterations that the unlocked version fails — and asserts the values match the mapped read, so removing the lock turns the suite red rather than the farm flaky Open
DC-124 An intermittent abnormal termination under memory pressure, seen twice and never reproduced Reproduced on the baseline, so it predates both threading changes. The two heavy gate tests — test_run_m1_gate (which runs the 4.20 GB checkpoint gate) and test_sharded_checkpoint — fail when they run adjacently, and pass alone: on e10a9ed without D101 or D102 the pair gave fail, pass, and with D102 it gave fail, pass as well. Same signature both times, so the cause is the harness's adjacency and this node's memory, not the fan-out. The reproduction is therefore: python3 -m unittest tools.test_run_m1_gate tools.test_sharded_checkpoint — roughly half the runs. The abort's own message is captured, which was this row's done-when. The line that looked like teardown noise is the fatal error itself, and the correlation is exact: it appears in 0 of every passing run and in every failing one (isolated sharded suite 0, the three passing m1-gate arms 0, each failing pair run 1–2, with the harness reporting -6). Its wording — "may have created a strong reference to self which outlived deinit, resulting in a dangling reference" — is the Swift runtime's resurrection-during-deinit diagnostic, so the mechanism is a thread messaging an UncachedFile whose last reference has just died. There are exactly two places one is held: Install.swift (private let blob) and Safetensors.swift (var file + streamingHandle()). Running the pair with NSZombieEnabled=YES MallocScribble=1 passed, which is consistent with a timing-sensitive race rather than a deterministic bug. Next diagnostic step, in order: a deterministic stress test that reads one install concurrently while copies of the source come and go. FOUND AND FIXED. Safetensors.StreamHandle.file was a bare var read and written from the reader's worker threads, so two workers could see nil together, each open a descriptor, and the racing assignments released one UncachedFile twice — the runtime's "deallocated with non-zero retain count … a dangling reference", then an abort. The tell was the discriminator: every failure was a checkpoint path and every install run was clean, because InstallFile holds its handle in a let and this cache did not. The handle is now opened once under a lock, with reads outside it. Evidence: the pair failed 3 of 4 runs before and 0 of 5 after, with the retain-count line appearing 0 times in all five (prior 75%; five clean runs by luck is ~0.1%). D101's round ran the full gates twice and got two different Python failures, both of which turned out to be the same event: a child process aborted mid-run. In test_run_m1_gate.py the gate's engine run died on the second prompt with the harness reporting "command terminated abnormally", swap at 1.21 of 2.05 GB used and only 3.61 GB reclaimable against the gate's declared 4.20 GB; in test_sharded_checkpoint.py it was exit -6 with the Swift runtime's "UncachedFile deallocated with non-zero retain count 2". That class is deinit { close(descriptor) } with no unowned/Unmanaged anywhere, so the line is very likely the runtime's teardown diagnostic for an abort that happened elsewhere, not its cause. Not reproduced after it was isolated: the m1 gate test passed 3/3 in a row (with the fan-out on, off, and on again), the sharded suite passed 5/5 alone, and datacenter-trace on the install passed 3/3. So it is not attributable to the expert fan-out, and the third full gate run is green. It is not explained either, which is why it is a row rather than a footnote Reproduced on demand with the abort's own message captured — tools/memory_watchdog.py running alongside, the child's stderr not swallowed by the gate — and either fixed or attributed with evidence Open
DC-123 Take the sister project's decode mechanisms — its int4 expert kernels, its residency handling and its prefetch ring — instead of writing them again, with the attribution that requires The operator authorised taking code on 2026-09-18, and D100 measured why it is worth doing: mix.read is 47.4% of a decode step, and inside it the CPU unpack is 40.5% of the step because the engine materialises 6.5 GB of fp32 per step and throws it away. The reference keeps expert weights in Metal buffers and reads a residency table from the kernel, which is exactly the cost this repository pays over and over. Apache-2.0 obliges attribution, not thanks: the first commit that takes a file lands the code, a root NOTICE (Copyright (c) 2026 André Borchert and the turbo-fieldfare dependency it names), the licence text, marks on modified files, and the check_provenance.py change that turns today's refusal into a requirement A file taken from TinyTitan is in the tree with its NOTICE and its marks, check_provenance.py requires that attribution instead of forbidding it, and the digest is unchanged Open
DC-122 Thread the contract matmul, or establish why it cannot be threaded unchanged The head's half is done (D102): the vocabulary blocks fan out and the phase fell 0.434/0.439 → 0.259/0.255 s/step with the digest identical in all four arms, because a whole row block is the unit — every block keeps the same Ops.orderedMatmul call and only the thread differs. HeadLogitsTests walks block widths 1, 3, 7, 64 and 512, deliberately including out % 4 != 0, which is where the general attempt moved its tail. What remains open is the general matmul: it is still the original single-threaded body, and its turn needs a decomposition that cannot regroup the four-wide columns — the menu for that is a case split per output, x-major blocking with a private accumulator, or accepting the four-wide grouping as the contract. DONE, both halves (D102, D104). The general matmul is threaded on width-aligned chunk boundaries, which is the rule the first attempt was missing: the serial body's four-wide groups are the fixed partition {0-3}, {4-7}, … and its scalar tail is exactly the last out % 4 columns, so a chunk that ends at a non-multiple of four re-cuts the grouping and moves the tail — which is what the reverted attempt did. With every boundary a multiple of four, a non-final chunk covers whole groups and has no tail, and the final chunk ends at out and performs the serial body's own tail. Measured on the phases, alternated: attn.core 0.493/0.495 → 0.209/0.202, mix.gateup 0.224/0.225 → 0.077/0.073, mix.down 0.102 → 0.049/0.047, step 4.124/4.103 → 1.866/1.797 s, digest identical in all four arms. OrderedMatmulThreadTests walks out % 4 ≠ 0, out < 4, k = 1, rows > 1, thread counts 0-16 and the model's real shapes. mix.read, head and attn.core are all Ops.orderedMatmul work on one core, and the parallelism looks free: every output is its own dot product over k, so splitting (row, column block) across threads changes which thread accumulates and not the order. An attempt was made and reverted (D99): it moved bits, and the repository's own reused-buffer comparison caught it — Ops.orderedMatmulVectorized is the original single-threaded body verbatim. The likely cause (a hypothesis, not a finding) is that bounding each block by its own last moves the four-wide grouping and the scalar tail relative to a single out-bounded loop for shapes where out % 4 != 0 or where the tail lands inside a block. DC-123's map threading shows the bar: identical digest across arms A threaded matmul whose output is bit-identical to orderedMatmulScalar on the contract grid, with the trace digest unchanged and the win measured on one binary — or a written reason why the contract order cannot be preserved across threads Open
DC-121 Prefetch the next token's experts while the current token computes And now the measured priority (D98): with capacity ruled out by the A/B above, hiding the reads is the only remaining way to stop paying for them, because the bytes must be read either way. Built, measured, and it LOSES — defaulted off (D105). The prediction is the reference's idea and the reasoning was sound: a layer's routing is correlated between consecutive tokens and the reads were on the critical path, so issuing them early on a background thread before the attention should hide them. Alternated on one binary: 1.859 / 1.816 s/step with it on against 1.854 / 1.721 with it off, and mix.read larger with it on (0.830 / 0.813 against 0.790 / 0.760) — the digest identical throughout. The reading is the finding: D101's fan-out had already removed the latency problem, so the device is now saturated and a second stream of reads does not fill a gap, it takes bandwidth from the first. That moves the lever from hiding the read to reading fewer bytes, which is DC-120's packed form. The reference's ExpertPrefetchRing stages predictions into raw-byte Metal slots and adopts them only if the router selects them; that plus its I/O event coordinator is where its 8-21% hidden I/O comes from. Ours hides 0% — there is no async file I/O in the engine at all. Worth doing after DC-118/119, not before: there is nothing to prefetch into while the bank dies with the layer A measured, non-zero hidden-I/O fraction and a lower host wait per token at the same hit rate Open
DC-120 Decode selected experts with an int4 GEMV from resident slots, never materialising fp32 Direction sharpened by D105. DC-121's measurement says the device is saturated, so the remaining win is not hiding reads but reading fewer bytes: the bank holds slices as [Float] today, so 512 MB buys about 61 of them while a token asks for ~773. The same bytes in their packed int4 form buy four times as many slices, which at 1.5 GB is roughly the whole per-token working set. That is now the target, and it needs the dequantise to happen at consumption — which is this row. D106 answered the caching half and the answer is memory, not bytes. A packed slab cache does cut the read — 45% hits, 1.08 → 0.65 GB/step at 1 GB — and the step still gets worse, because load and head grow more than mix.read falls on a swapping machine. So fusing the dequantise into the matmul remains worth doing for the unpack and the fp32 materialisation, but a bigger cache is closed as a route on this node. CLOSED as a CPU route (D107), on a measurement that also bounds it. The fused int4 matmul was built and proved bit-identical to dequantise-then-multiply over a grid of shapes and the model's real ones (512×2048, 4096×2048) — so its speed was the only open question, and the answer is 4.7× SLOWER in release: 0.33 ms for the split path against 1.55 ms fused (0.14× in debug). Dequantising inside the inner loop destroys the vectorisation that dequantizeInt4's eight-wide path and the matmul's four-wide path each enjoy. The probe also bounded the prize: a 512×2048 slab unpacks in 0.23 ms, so all ~520 slabs a token cost about 0.12 s/step, ~7% — the fused form had almost nothing to win even had it been fast. Reverted, not kept, on the DC-122 precedent: the number is worth more than the code. The reference keeps expert weights in Metal buffers and reads the residency table from the kernel (MetalExpertReader, MetalExpertStagingPool, Kernels/MoE/). Ours reads bytes, dequantises to an fp32 array on the CPU, then either matmuls on one core or (by default) is not on the GPU at all — MetalMatmul is opt-in by a verdict that predates its tiled kernel The expert phases run on the GPU with the accumulation order unchanged, verified by the existing 288-shape bit-identity test and the trace digest Open
DC-119 Let the expert bank outlive the layer, and instrument it Built and measured 2026-09-18 (D98): the lifetime fix works and it answered the question the other way round. ExpertBank holds slices for the whole generation, byte-budgeted and LRU, with the per-projection cap and peak the old metric meant, and it reports hits, misses, elements read and resident peak into metrics.json — none of which existed. On the real 35 B-A3B, 24-step decode, five alternated runs: 0.0% hits at 0, 512 and 1024 MB, with identical 29,192.4M elements read, and the bank on was slower in both pairs (4.020 / 4.132 s/step off, 4.139 / 4.236 on; 4.335 at 1 GB). The arithmetic says why: 773 requests per step, 12.5 MB per slice, so one token's working set is 387 slices per projection = 4.83 GB, ~9.66 GB for both — a reuse distance nine times a 537 MB bank. Capacity cannot be the lever on an 8 GB node, so the default is 0 (as D89 did for the layer cache) and the store is what the prefetch ring will stage into. D31's 0 was right about the workload too; the lifetime fix is what made the two distinguishable. 233 Swift tests, 7 new. case mixture(_, provider: ExpertSlotCache) is built inside loadLayer and dropped with the layer, so on a cached decode the hit rate is 0 by construction — D31's measurement was right and the cause was the lifetime, not the size. The reference gets 60-85% from a bank that lives for the generation, and its curve is worth +37% from 1 to 3 GB. Ours also reports no expert metrics at all: not hits, not bytes, not read seconds A generation-scoped (layer, expert) bank sized by SHARD_EXPERT_BANK_MB, with hit rate, bytes read and read seconds in metrics.json, and the trace digest unchanged Open
DC-118 Fan expert misses across threads, as the reference's streamer does Built and measured 2026-09-18 (D101): the fan-out is worth 1.35x. ExpertWeightProvider.preload(experts:shape:) is a hint the mixture gives once per layer, and the adapter fans the misses across threads into the bank as staging — deliberately not counting them as requests, because the loop that follows is the requester and counts a hit. Alternated, one binary, 8-step decode: 3.383/3.382 → 2.503 s/step (mix.read 1.671 → 0.761), digest identical in all four arms; the bank without the fan-out is 4.109 s/step, so the pair is the configuration and the default is now 512 MB. Instrument caveat: with concurrent reads the per-thread counters sum past wall time (read 3.867 s and unpack 5.206 s in a 2.503 s step), so they measure thread-time under overlap, not elapsed time. The reader's counters are write-locked for the same reason. PreadExpertStreamer reads its misses with DispatchQueue.concurrentPerform; ours issues one direct read at a time. The reads are independent, so this is the D94 shape again — the same trick that took load from 1.33 to 0.56 s/step — and the bytes cannot change a value (docs/reference-tinytitan-decode.md) A decode step's mix.read falls with no change to the trace digest Open
DC-117 Reach 7 tok/s decode on one node — the operator's precondition for any further network work Progress 2026-09-18 (D99): the element-wise passes are threaded and the step fell 4.119/4.162 → 3.448/3.465 s/step in an alternated A/B on one binary (load 1.097 → 0.356), i.e. 0.242 → 0.290 tok/s, with the trace digest identical in all four runs. mix.read is now the top phase at 1.71-1.77 s (50%). Distance to 7 tok/s: ~24x. Progress 2026-09-18 (D101): the expert reads are fanned across threads and the step fell 3.383 → 2.503 s — 0.400 tok/s, from 0.230 when this row was opened. Progress 2026-09-18 (D102): the head's blocks fan out too — 0.443 tok/s, from 0.230 when this row was opened. Progress 2026-09-18 (D104): the contract matmul is threaded too — the step is 1.797-1.866 s, about 0.55 tok/s, from 0.230 when this row was opened. The engine measures 0.230 tok/s (4.32 s/step) for the 35 B-A3B on a Mac mini M2, and the sister project measures 7.075 tok/s for a 35 B-A3B at 4-bit on a comparable 8 GB M-series node — a ~30x gap that is architectural, not tuning. Five differences, in the order they cost time: (1) the expert cache is born and dies inside one layer load (case mixture(_, provider: ExpertSlotCache) is built in loadLayer and dropped with the layer), so cross-token reuse is impossible and the hit rate is structurally 0 — while the reference gets 60-85% from a bank that lives for the generation, and its curve (5.164 → 6.019 → 7.075 tok/s at 1 → 3 GB) says what that is worth; (2) the GPU is off — MetalMatmul is opt-in by D63's verdict, which predates the tiled kernel now in the source, so the reference's 33-46% GPU busy has no counterpart here; (3) no I/O is hidden — there is no async file I/O anywhere in the engine (no DispatchIO, no reader thread), so 0% of expert reads overlap compute against the reference's 8-21%; (4) the per-token matmul is one core with a k-strided access pattern (~1/16 of usable bandwidth: four output rows 8 KB apart per k step); (5) the head materialises 2.03 GB of fp32 from bf16 per token, single-threaded, for 1.05 s of a 4.32 s step. Progress 2026-09-18 (D108): the fused GPU int4 matmul is built and bit-identical, and it is worth 0.114 s of the 1.74 s step — about 6.6% — which says the arithmetic is not the lever: mix.read is the disk read, and the distance is closed by residency, not by a kernel. Progress 2026-09-18 (D109): the LM head now runs on the GPU from its stored bf16 and head fell 255 → 118 ms; the default step is 1.567 s = 0.638 tok/s with the digest unchanged. Three levers were measured and refused (dense GPU unpack; the fused int4 expert path, a wash at 1.575 s against a 1.69 s control; head residency until two allocation bugs were fixed), and the study of the sister project says the remaining distance is batching the experts per layer into one dispatch, then residency, then overlap — its own 141 ms is ~55 ms of it host-blocked on the routing readback and the miss reads. Progress 2026-09-18 (D110): the int4 kernel was loading one byte per element at 4.4 GB/s; reading each row as a uint4 made it 3.6-4x faster, and with a 256 MiB packed slab cache (larger sizes lose to memory pressure) the fused expert path is 1.465 s/step against the split path's 1.539 over five alternated pairs, digest unchanged. The default configuration is now 0.682 tok/s A single node reports ≥7 tok/s decode for the 35 B-A3B with the trace digest unchanged, measured with conditions alternated on one binary and the load recorded — and only then does cluster measurement resume (D97) In progress
DC-116 Establish whether dense models belong to this design, and record the answer The target was an overview across Qwen3.5 4B, 9B and Qwen3.6 35B-A3B, so a 4B install was built from the real checkpoint (426 tensors, 4,204,789,760 weights, 3.34 GB at 6.36 bits/weight) and measured on one node. Dense is out, and the measurement says why: the shard plan divides experts, so a dense model has no plan at all and cannot be distributed; and its whole payload is re-read for every token — 18,638,208,000 bytes in four steps, with 0 hits, because the default 1 GiB SHARD_DENSE_CACHE_MB cannot hold a 3.34 GB payload. Decode came out 5.870 s/step = 0.170 tok/s, i.e. slower than the 35 B MoE on the same node (0.230 tok/s), which activates only ~3 B of its 35 B. The 9 B is worse still: 5.5 GB of payload against ~4.5 GB of usable RAM, and still undividable. The table carries the 4 B row with its prefill unmeasured, and the README says dense is outside the design The answer is recorded with its number, D96 carries the reasoning, and the README's table shows the MoE rows beside the dense one rather than leaving a reader to wonder why 4 B and 9 B are absent Done
DC-115 Cut the first release (v1.0.0), with an identity a build can enforce and an archive a user can verify Done and verified 2026-09-17: v1.0.0 is published with the archive and its checksum, the notes quote digest b6ebfd1b861a1ac4… and the .sha256 beside the archive matches it, shasum -c passes on the downloaded bytes, and all three shipped binaries answer 1.0.0 and report arm64. One blemish, on the record rather than rewritten: the archive also carries two test bundles, because the first version carried every bundle the build directory held instead of the set the plan says the executables can reach — harmless, 2 MB of fixtures in bin/, and wrong. The selection is now plan-driven (declared_bundles, with tests) and reports (none declared) for this package; the published assets are left alone, because rewriting a tag someone may already have downloaded is worse than a blemish that is written down. RELEASE.md §1.3 wants the version single-sourced and enforced, and §1.6 wants the archive to carry what a Swift binary needs. There was no VERSION, no --version, no release script and no changelog, so the identity had to exist before a release could. tools/release.py runs the gates, does a clean scratch build with the log scanned for warnings, packages the three executables with the licence, the notices and a README-binaries.txt, asserts lipo -archs on the binaries extracted from the archive, and refuses to publish notes that do not quote the digest it just computed (D95) A published release whose assets are the archive and its checksum, with --version answering 1.0.0 from the binaries inside it, and tools/version.py --check refusing a drifted mirror In progress
DC-114 Overlap the exchange wait with the next layer's decode Not needed for the objective 2026-09-17 (D94): the cluster reached 1.74x without overlapping the wait, so this row's done-when is now about the roadmap's ≥3x rather than 1.5x. The cluster is 4.021 s/step and needs 3.658 for 1.5x, and its largest single phase is load at 1.9 s/step — 47.9% — during which the GPU idles and the CPU dequantises constants (D88). A node waiting for a peer's contributions is idle too, and the work it will need next is known: decoding the following layer's weights during the wait turns one into the other. It must be one layer of decoded weights held transiently, not a whole-model cache: D89 measured that holding 16 layers loses to memory on an 8 GB node, and the difference is that a one-layer prefetch holds bytes the next layer needs anyway. A cluster step at or below 3.658 s/step with bit-identity intact, or a measurement showing the wait is too small to hide anything in Open
DC-113 GPU decode: a quantized GEMV that dequantises in-kernel Decode is M=1, so the right GPU kernel is a GEMV that reads int4 weights and dequantises inside it. D88 measured load at 30.5% of a cached step — the same ~1 G parameters of constants dequantised on the CPU on every token — and D89 showed caching them loses to memory on an 8 GB node, so the phase has to stop existing rather than be paid for twice. It can stay bit-exact (D63: one thread per output, k accumulated in order — tiling changes which thread works, not the order of the sum), which an ANE path cannot (D91). Any device selection must verify the device it got and refuse a silent CPU fallback: the sister project's measured trap is Core ML exiting 0 while running on the CPU at ~38x the GPU cost (D91). Superseded on the path to this number 2026-09-17 (D94): the CPU dequantiser's row loop now runs across the machine's cores, which took the same phase down 2.4x (1.33 → 0.56 s/step) with no device, no protocol, no bit-exactness argument to make and none of the silent-CPU-fallback risk D91 found. The GEMV remains a candidate for absolute speed, not for this objective. Measured, and blocked by memory (D100): SHARD_GPU_MATMUL=1 was run and the disk watchdog stopped it at 3.94 GB free with swap at 1.71 GB — the head's 1.017 GB of bf16 becomes 2.03 GB of fp32 and the kernel needs an equal MTLBuffer, about 4 GB on a node with ~4.5 GB usable. The tiled kernel's speed is still not measured; what is measured is that the fp32 array has to go first (DC-120). Now the route, not a candidate (D107). With the read saturated (D105), the cache unaffordable (D106) and the fused CPU kernel 4.7× slower with only ~7% at stake, every CPU-side lever is measured and closed: the engine's split of dequantise-then-multiply is the right shape for a CPU. The distance to 7 tok/s is the format and the device — packed int4 weights resident in GPU buffers, dequantised in the kernel, with the residency table and the prefetch ring — which is this row. Built and measured (D108): the dequantise-in-kernel half is real and it is not the lever. MetalInt4Matmul computes x @ dequantizeInt4(w)ᵀ in one dispatch, bit-identical to dequantizeInt4 followed by Ops.orderedMatmul over the D107 grid, and 1.23x / 1.45x / 2.36x faster than unpack-then-matmul on the real shapes — worth 0.114 s of a 1.74 s step, about 6.6%, the same order as the 7% D107 bounded from the CPU. mix.read is the disk read, so the 7 tok/s route is the residency half (the residency table and the prefetch ring), which is DC-127; this kernel is that work's precondition, not its win. The grid also found a boundary that had been assumed away: an Apple GPU flushes a denormal product to zero where the CPU keeps it, for this kernel and for the MetalMatmul already in the tree, and pinned it in MetalInt4MatmulTests rather than leaving it implicit. — The decode step's matmuls run on the GPU from quantized weights, bit-identically, and the step is measurably faster Open
DC-008 ADR — transport: Thunderbolt bridge vs LAN vs SFP/QSFP, measured latency and bandwidth per hop. The implementation half is done (D22): a listener that binds, listens and accepts, a connector with an explicit deadline, and a contribution exchange over TCP that is bit-identical to the single-node forward. What remains is the half the ADR is named for — the per-hop measurement — which needs the testbed, plus IPv6 if the farm needs it DC-009 The tabled latency and bandwidth come from measurements on the real links, not from the paper In progress

DC-004 status: CI carries the repository's own gates on every push — the Markdown link gate (tools/check_markdown_links.py: local links and #anchors, offline), the table gate (tools/check_markdown_tables.py, over the repository and this wiki) and the whole tools suite. The Swift workflow (.github/workflows/swift.yml) fails on a runner that cannot meet the project toolchain instead of warning and skipping (DC-103): GitHub's macos-26 image carries Xcode 26.0–26.5, so that job is red by design until an image ships Xcode 27, and the Swift gate runs locally until then. The tools runner has no torch, so the capture tests skip there.

Closed in this phase: DC-001 Bootstrap the wiki: Home, Roadmap, Project Tracker, · DC-002 README correction pass: the Shard codename in the · DC-003 Repository scaffolding: .gitignore, AGENTS.md, · DC-006 ADR · DC-007 Golden-trace harness specification: trace schema, · DC-010 ADR · DC-012 Repository layout and SwiftPM target plan, consistent · DC-014 Language standard recorded and applied: Swift 6.4, · DC-015 The Swift language-feature register: which upcoming · DC-016 Core ML tooling installed and verified: coremltools · DC-017 Farm inventoried from this checkout: every node's · DC-019 Interconnect, storage and memory baselines measured, · DC-037 M1's numeric contract: the mixture of experts · DC-038 Expand the key and query heads up to the value head. Evidence on News.

P1 — M0: staged, dense Qwen3.5-2B and its harness

The harness is proven before the model, and the model is proven in the order that keeps one variable at a time (D1):

Closed in this phase: DC-020 Choose and pin the M0 model: Qwen3.5-2B (Apache-2.0) · DC-021 Reference runner that emits golden traces for the · DC-022 IR importer for the M0 models (name-to-role mapping · DC-023 Minimal single-node forward pass, driven by the IR, · DC-024 Trace-capture harness (per-layer activations, logits, · DC-025 Diff harness with bit-exact comparison and · DC-026 Gate M0: on the pinned model, zero differing bytes · DC-027 Reference contracts for both M0 models: the official · DC-028 M0a: capture and diff harness, proven able to · DC-029 M0b: Qwen3.5-2B bf16 bit-exact, layer by layer,. Evidence on News.

P2 — M1: Qwen3.6-35B-A3B, single node, 4-bit, SSD-streamed

ID Task Depends on Done when Status

Closed in this phase: DC-030 Importer for the Qwen3.6-35B-A3B · DC-031 Quantization/transcode · DC-032 SSD expert streaming with a bounded per-node expert. Evidence on News.

Sweep result (2026-09-16). capital's five tokens at bank sizes 2, 8 and 16 produced three byte-identical traces — trace_diff reports 83 tensors, 0 differing elements and 40 discrete decisions — with identical expert counters at every size: 6,977,224,704 bytes from SSD (1,395 MB per token), 2,218 requests, 0 hits, 13.95 GB decoded in memory, and 39.3–39.9 s of prefill (7.90 s per token, 0.127 prefill tok/s — a prefill rate, not the gate's generation baseline). The digest moved from the pre-D11 b8c976c5… to b0d382db… exactly as D11 predicted. Full numbers and the one unexplained figure (2,218 requests against 320 selections per token) are in docs/m1-gate.md.

| DC-107 |The kernels are done and they are the wrong lever — the engine is I/O-bound (D64). The tiled kernel is bit-exact (288-shape grid, head shape, cache reuse; the real trace b0d382db… either way, trace_diff IDENTICAL) and still slower than the CPU: attn.core 4.05 → 5.6-5.9 s, mix.gateup 1.18 → 2.1 s. Coalescing the loads changed nothing, so the cost is the path around the kernel — an encoder, two copies and a synchronous wait per call, ~4 ms across ~210 calls in mix.gateup — not bandwidth. And the profile says where a forward goes: mix.read + load + head are 12.73 s of 18.75 s, 68%, at the ~1 GB/s sequential floor. So the lever for M3's ≥3× is sharding the reads (M2/M3 already do), not arithmetic. The kernels stay opt-in (SHARD_GPU_MATMUL=1) as a tested asset. Remaining targets here: the reads, and attn.core's 4 s of compute, both bounded by the same 68%. The GPU matmul is bit-exact and SLOWER, so it is opt-in (D63) — and the head's 1.74 → 1.28 s was drift. Routing every contract matmul (38 call sites) through the chooser leaves correctness untouched — the trace is b0d382db… on the GPU and on the CPU, trace_diff IDENTICAL — but with the conditions alternated rather than run in sequence, every phase that uses it is slower: attn.core 4.05 → 5.55 s, mix.gateup 1.17 → 1.62 s, head 1.74 → 1.86 s, while mix.read (no matmul) is unchanged, so it is the kernel and not the machine. Cause: one thread per output walks w rows k*4 bytes apart, so a warp touches 32 cache lines for 4 useful bytes each; the CPU's path walks k contiguously. Fix: a threadgroup-tiled kernel, which cannot change a result because tiling changes which thread accumulates, not the order. SHARD_GPU_MATMUL=1 until then. The head still reads 1.02 GB of its own weights, which is its floor. The head is done and it was half compute, half I/O (round 45). The GPU matmul — one thread per output, ascending k, acc + metal::fma(x, w, 0), the spelling D61 measured — is bit-identical to Ops.orderedMatmul over 288 shapes (including output counts that are not multiples of four, where the CPU vector takes its tail), on the head's real shape and across cache reuse; the real trace is b0d382db… with it on and off, trace_diff IDENTICAL. Measured by phase: head 1.74 s → 1.28 s, while the totals said nothing (16.3 against 16.4 s, the wrong way round, because the reads vary more than the saving). The rest of the head is 1.02 GB of head weights at ~0.8 GB/s — the same sequential floor as the expert reads — so it is documented as at its limit. attn.core 4.05 s did not move and is the opposite case: small activations, arithmetic-dominated, and the right place for this kernel. The head kernel's arithmetic is now settled (D61). Before writing it, the question D10 left open was measured: .fast, .relaxed and .safe all contract a * b + acc into one fma (.safe only keeps a product in its own statement apart), so no mode is enough — a GEMM must write accumulator + metal::fma(x, w, 0.0f), which reproduces the contract's separate rounding under all three modes and lets it share .relaxed with the unpack. The probe uses a one-ULP discriminator for that reason: the same probe over a 257-term dot passed in every combination, so a long-dot test would have declared the GPU exact and been wrong. Re-profiled (round 43), and the order changed: mix.read 7.23 s of a 19 s trace — the install read 3.24 s plus the unpack 4.53 s — then attn.core 4.05 s, load 2.22 s, head 1.76 s, mix.gateup 1.22 s, mix.down 0.53 s. The unpack is no longer a target: the GPU path is now the default, verified bit-identical three ways, and takes about two seconds off the trace (D59, D60) — so the earlier note that it was "at its measured limit" was wrong and is corrected here. The read genuinely is at its limit: 3.24 s for 3.06 GB is ~0.95 GB/s, the sequential floor, and no kernel changes that. The next kernel should be the head, a plain GEMM (1.76 s) and the easiest of the remaining ones to make bit-exact, before the DeltaNet's chunked rule inside attn.core. The unpack target is met (round 42). The GPU unpack is now faster than the scalar one with the same digest: 16.3 s against 38.2 s before the buffer cache and 18.2 s on the CPU, with trace_diff against the scalar run reporting IDENTICAL — 83 tensors, 0 differing elements, 40 discrete decisions. The fix was the plumbing the diagnosis named: a one-slot buffer cache held under a lock across the dispatch and its completion, and one copy per section instead of four. Recorded as wall-clock observations of one trace each, not a benchmark; the default stays opt-in until the timing phase, when the frozen prompt set can be the evidence (D59). The remaining targets here are attn.core and the reads. Round 41 measured the unpack's GPU path: same digest, slower — 18.2 s scalar against 38.2 s GPU for the five-token trace. The kernel is not the cost: the row path concatenates codes + scales + zeros and unpack copies it again, and a fresh MTLBuffer is allocated per expert fetch (2,218 of them). So the remaining work here is the plumbing — a reusable buffer, or a dispatch over a payload already contiguous — not the shader. The next performance targets: attn.core 3.86 s, the expert unpack 3.45 s, the reads 2.99 s. After D15 and D16 removed 24 s of hashing and duplicate reads, these are what a 14.97 s five-token forward is made of — none of them I/O, and the reads and unpack are at their measured limits (~1 GB/s and ~1,012 M values/s). The attention core is the largest single compute phase and the first target | DC-105 | Each is either faster with the same digest, or documented as at its limit with the measurement that says so | Planned | | DC-112 |Decided (round 40): the DeltaNet's and the attention's projections should return to bf16. The logits settle it: against a bf16 checkpoint contract the argmax is unchanged on all five positions, but the smallest margin is 0.27, top-8 overlap is only 5–8 of 8, and the mean logit difference is 0.30. That is the same situation the policy already treats as decisive for the router and the gates — they decide the discrete outcomes — and a narrow pass is not a reason to spend it. The routed expert stacks (~19 GB of the 21.7 GB install) stay int4, which is where the size win lives. Execution is blocked here: the rebuild needs ~25 GB free and this node has 9 GB, and it is a GB-scale job that must not run on this machine. The gate conflict is settled (round 39); what remains is a quality-and-size choice, not a correctness one. The milestone's claim is restated to the form that can be falsified and proved: the engine's trace is byte-identical to a contract reading the same install (IDENTICAL — 83 tensors, 0 elements, 40 discrete decisions, matching digests), verified 2026-09-17. The design decision left open is whether the DeltaNet's and the attention's projections stay int4 — about 3 GB of a 21.7 GB install, with the routed expert stacks' ~19 GB unchanged either way — or return to bf16 to narrow the gap to a checkpoint contract, now measured at 0.0156 max / 24.5% median relative at layer 0. A rebuild to test that needs ~25 GB free and must not run on this node. Round 39, the way to settle it without a rebuild: give both sides the same weights. D55 showed the whole divergence was weights, not arithmetic, so the gate's comparison had two inputs and could not falsify anything. tools/install_source.py feeds the contract from the install through the same dequantiser the Swift reader mirrors, so a disagreement is a disagreement in arithmetic. Findings: an install flattens trailing dimensions, so one expert is 1,024 consecutive install rows of the 262,144-row expert.stack_gate_up — D32's "one expert is one row" is true of the checkpoint, not the install; Install.row_range makes expert 255 cost one expert; ten new tests (359 total). The full-model run is in flight — 1,600 expert fetches through Python, 878 MB and ~19.5 min CPU at last measurement, both guards live. Not claimed until measured. Settle the conflict between M1's gate and the quantisation policy (raised by D55). Either keep the DeltaNet's and the attention's projections at bf16 — roughly +3 GB on a 21.7 GB install, since the routed expert stacks are about 19 GB of it and stay int4 — or restate M1's claim with the measured cost of int4 as its evidence: 0.0156 max / 24.5% median relative at layer 0. Rule 3 forbids the quiet version of either. Blocked on resources here: the rebuild needs about 25 GB free and this node has 9.5 GB, and it is a GB-scale job that must not run on this machine at all | DC-111 | The decision is recorded, and either the policy or the claim changes with the measurement beside it | Open |

P3 — M2: 2 nodes, expert-parallel

ID Task Depends on Done when Status

P4 — M3: 4 nodes

ID Task Depends on Done when Status

| DC-051 | Per-token synchronisation budget: all-reduce latency per MoE layer against expert-read time Measured 2026-09-17 (D84): the budget is 0.893 s/step of exchange — 15.0% of a 5.975 s step — moving 4.49 MB at 5.0 MB/s effective, i.e. 1.63 ms per exchange term over 547 terms/step and 90 reduces/step, which is a per-message latency rather than a saturated link. That accounts for the exchange; it does not account for the ~3.8 s/step that is neither payload read across the plan nor exchange, because the node reads 2.65x less than the baseline and is slower. Remaining: a per-phase breakdown of that 3.8 s, which is the instrument this row and DC-107 still need. Step-level budget measured 2026-09-17 (D88), on the cached decode the gate measures: mix.read 1.80 s/step of device reads at 184 MB/s (0.33 GB/step), load 1.66 s/step of CPU dequantisation of constants, head 1.05 s/step of matmul on cached weights, attn.core 0.52 s/step. The exchange (D84) is 0.89 s of the cluster's step. The plan therefore addresses 33% of a step, and the target is not reachable from it. The cluster's own budget measured 2026-09-17 (D90), per node: the exchange is 4.250 / 4.136 / 0.672 / 3.739 s/step on nodes 0-3 — 71.5%, 69.5%, 11.2%, 63.1% of the step — while the plan is demonstrably working (reads 6.281 → 2.31-2.40 GB per node, mix.read 1.80 → ~0.49 s/step). A 6× spread between nodes doing identical work is the signature of head-of-line blocking: allReduce sends to every peer and then receives sequentially in peer order, blocking, so one slow peer delays everyone behind it, forty times per token (~105 ms per layer). The fix is the budget's other half: receive concurrently and skip peers that own none of the chosen experts. And the counter does not reconcile with the phases (2026-09-17, D93): the ledger's exchange_seconds is 1.4-1.9 s/step on every node while the ff phase that contains it is 0.575 s, even though ff is marked after mixtureOutput returns and the phases sum to the step (3.9 of 4.02 s). Both cannot be right, and D90/D92 reasoned from that counter, so it is recorded rather than used until it is resolved. | DC-045 | The measured budget explains the achieved scaling | Planned | In progress | DC-052 | Per-node cache and prefetch tuning at four nodes Target identified by measurement 2026-09-17 (D88): the cached step's load phase is 30.5% of a 5.449 s step — 1.66 s/step — and it is not I/O. The dense payload is read once for the whole run (1.0437 GB held, 4888 cache hits), so that time is the same ~1 G parameters of constants being dequantised and released on every token, 160 times per generation, on the CPU while the GPU idles. mix.read (33.0%, 1.80 s/step, 184 MB/s) is the only device-bound phase and the only one a plan divides. So this row is the speed lever: decode once and hold, with the budget measured and recorded per node rather than guessed. Built and measured 2026-09-17 (D89): the decoded-layer cache holds layers across the steps of a generation — not an LRU, which a sequential sweep gives a zero hit rate by construction — counts the bytes it holds, and records budget/held/hits/misses per node in metrics.json. It is bit-identical (asserted tensor by tensor) and it works: about 0.035 s per layer per step, 1.4 s across all 40. On this 8 GB node it still loses — at 1 GB and 2 GB the step rose to 5.46 and 5.60 s from 5.33 s, with head +0.20 s/step and attn.core +0.27 on a machine already swapping — so the default budget is 0, and SHARD_LAYER_CACHE_MB is the knob for a node with headroom. Prefetch depth remains open. | DC-050 | Cache budget and prefetch depth are recorded per node | Planned | | DC-053 |Round 59: the INSTALL form of M1's gate also passes, all five frozen prompts. DC-114 was fixed — the install dequantiser is vectorised, bit-identically, 0.253 s to 0.013 s for one real expert — and the gate that could not finish a five-token prompt in twenty-five minutes now completes the whole set: capital 17.40 s, arithmetic 54.51 s, code 67.63 s, repeat 86.90 s, long 98.45 s, every one IDENTICAL — 83 tensors, 0 elements, 40 discrete decisions — GATE PASSED, digests b0d382dbabf36df0… on both sides, peaks 0.98-1.41 GB. That is M1's restated claim (D56) checked by its own instrument, alongside the checkpoint form checked in round 56 (D73). The install declaration moved 1.5 -> 2.0 GB because the measured peak is 1.41 GB and a 6% margin is a coincidence rather than a guard. Round 56: M1's gate PASSED on the real checkpoint, on this node, all five frozen prompts. The checkpoint path was still going through safe_open — mapping 67 GB, the hazard the node's rules name — because the gate passed only one of the two flags docs/m1-gate.md records as the survivable pair; it now passes --uncached on the checkpoint path too. Measured first: the contract on the checkpoint is 66.20 s and 0.397 GB peak, against 4.16 GB recorded before the flags. Then the gate: capital 34.94 s / arithmetic 89.65 s / code 115.92 s / repeat 128.13 s / long 146.26 s, all IDENTICAL — 83 tensors, 0 elements, 40 discrete decisions — GATE PASSED, peaks 3.60-3.79 GB, and the first prompt's digest b8c976c5e7ba8816… is the one docs/m1-gate.md recorded on 2026-09-16. The declaration for this path is 4.2 GB, restored from a 1.5 GB I set from the contract's measurement until the gate's own 3.79 GB peak corrected me: a checkpoint run's peak is the engine's trace, not the contract's (D73). What remains here is the install path (DC-114) and the ≥3× measurement in a quiet window. Round 51: the four-node mesh is re-verified and the functional half is runnable on its own. With the farm quiet (node1 2.5, node2 1.4, node3 1.0) run_m3_gate.py --functional-only ran the full mesh across four machines against the single-node run: every node IDENTICAL, tokens and digests, over a plan staged to three peers that already held their own installs. The new flag expresses the operator's own split — it checks the cluster and stops, computing, printing and recording no speedup, and nulling the timing fields rather than omitting them so the output cannot be mistaken for a measurement; it also skips the quiet-farm precondition and says so, because a functional run makes no throughput claim. What remains here is the ≥3× measurement in a quiet window, which is deliberately not run yet. Round 48: the tokens are verified and now gate-checked; the gate itself still needs a bigger node. The cached and full-sequence paths generate the identical eight tokens (11751,11,264,3177,34756,364,1141,8807) with identical top-2 margins at every step — including 0.1339 and 0.0445, which are narrow enough that a numerical difference could move them — and run_m1_gate.py now runs both and requires equality (I3, no tolerance), with an empty parse failing rather than passing. The gate's own fixture test exercises the new path. What remains for this task is unchanged and is the throughput measurement in a quiet window; Corrected in round 52: run_m1_gate.py now discovers whether it was handed a checkpoint or an install and runs both halves against it, and it passes --stream-experts to the contract — without which the contract half could not read an install's expert stacks at all (D69). The install form needs no 70 GB mapping, so M1's own gate is runnable on this node; what D65 said about the checkpoint form still holds. Gate M3: ≥3x the frozen M1 tok/s on the same prompt set. The functional half is demonstrated (D27: four machines in a full mesh, one bit-identical trace) and the instrument now exists (D28: datacenter-generate shards, and a two-machine cached generation produced the reference's tokens and digest). What the gate needs is the measurement, on a quiet farm, since the nodes are shared with other work First real measurement 2026-09-17, on a busy farm with the operator's authorisation (--allow-busy-farm, so the gate recorded it as observation_only with the loads and did not assert the threshold): bit-identity holds on all four nodes, and the ratio is 0.93x — baseline 5.662 s/step against the slowest node's 6.086 s/step. The per-node metrics show why: a node reads 2.65x less payload per step (0.592 GB against 1.570 GB, because the dense 0.261 GB is replicated and the experts really are 3.95x fewer) and is nevertheless slower, and the exchange is 0.893 s/step (15.0%) moving 4.49 MB at 5.0 MB/s effective — 1.63 ms per term, latency-bound, not bandwidth-bound. So the >=3x target is not what an expert plan delivers on this design: the step is not expert-read-bound. The threshold remains unasserted and this is not a gate result — four steps, a busy farm, and the quiet-window measurement is still owed (D84). And measured 2026-09-17 (D94): the cluster runs at 1.74x and 1.70x over a single node with eight decode threads, and at 1.66x with one — four alternated runs, every one bit-identical on all four nodes, loads 1.6-5.5. The operator's objective (≥1.5x) is therefore met on a busy farm, and load hurts the ratio (four nodes are exposed to a spike and the baseline is one), so these are lower bounds. The roadmap's own ≥3x on a quiet farm is still not asserted, and the gate remains the certification instrument. | DC-045 | On a quiet farm, the same prompt set runs at ≥3x the M1 rate with bit-identity intact | In progress |

P5 — M4: DeepSeek-V4.1-Flash

ID Task Depends on Done when Status
DC-060 DeepSeek importer with transcoding of native FP4 experts and FP8 dense weights DC-045 No requantization step exists in the path Planned
DC-061 Compressed / sparse hybrid attention and constant-size recurrent state DC-060 The model runs at the reference context length Planned
DC-062 MTP speculative decoding parity DC-060 Draft and verify match the reference protocol Planned
DC-063 Sparse-attention block-selection exactness check DC-061 Selections match the reference exactly, not approximately Planned
DC-064 Gate M4: matches the reference at 128K context DC-063 The exactness protocol passes at 128K Planned

P6 — M5: Qwen3.8-Flash-Next

ID Task Depends on Done when Status
DC-070 Importer and reference match for Qwen3.8-Flash-Next DC-064 Same exactness protocol as M4 Planned
DC-071 Sparse n-gram lookup table handling and per-node sharing DC-070 Table lookups are correct and bounded in memory per node Planned
DC-072 Expert-parallel scaling on the 125B-A6B shape DC-070 Gate M3's ratio is reproduced or explained on this shape Planned
DC-073 Gate M5: matches the reference, same protocol as M4 DC-072 Recorded, with the measurement Planned

Cross-cutting

ID Task Depends on Done when Status
DC-085 Re-audited 2026-09-17 against the same snapshot, and the position holds with evidence. Its int4 is unsigned with a bias, and its own reference says so in the first line of the function that consumes the format: its own qwen35_reference.py's dequantize() — "bits-wide unsigned lanes packed low-first inside each uint32, one BF16 scale and bias per group" — computing grouped * scales + biases, with dequantizeInt4Affine on the Swift side. Ours is signed codes with an int8 zero point (_signed() and (code - zero) * scale in tools/install_reader.py), and its container family is GTurbo*V1 (GTurboFormatV1, GTurboExpertV1, GTurboLayerV1, GTurboManifest*V1) against our install.json plus data.bin. The same bytes mean different numbers under the two rules, so a file from one is not readable by the other — arithmetic rather than opinion. No defect to report upstream, because its reader and converter agree with each other. The checkout carries no git history, so "track its releases" cannot be done from it; the snapshot is a source drop dated before the original audit, and the position is recorded with its evidence in THIRD_PARTY_NOTICES.md (D80). Audited 2026-09-16: the sister project's int4 is unsigned-with-bias and its container family is GTurbo*V1, so there is no defect to report and nothing of ours was copied (D39). Synchronisation with the sister project: track its releases, report cluster-relevant defects upstream — A defect found here has a linked upstream issue where it belongs there Planned

Closed in this phase: DC-082 Documentation sync rule: README, wiki and this · DC-086 Uncached reads for model payloads: open the payload · DC-088 Per-slab digests, so a reader can verify the expert · DC-089 D11 option 1: flush denormal scales to zero in both · DC-093 WITHDRAWN · DC-094 The spec/config correspondence is a tool with tests · DC-095 The wiki's tables have never been gated in CI, and · DC-096 Audited the two largest unprobed engine surfaces and · DC-097 The brief's 16 KB alignment is unimplemented, and · DC-098 I6's two gaps are closed, all three items verified. · DC-099 Audited I4 and L2 and found both · DC-100 The M1 sweep is a script that refuses to start, not a · DC-101 The brief's own · DC-102 The public status claim is corrected in all five · DC-103 The toolchain is enforced, not documented: Xcode 27 · DC-104 Corrected the M0 status claim in six documents: M0 is. Evidence on News.

Risks and traps

ID Risk Handling
R1 Bit-identity may not be attainable on Metal. Nondeterministic reduction order or driver scheduling can make two runs differ. The deterministic reduction contract (DC-005) is written before M0, and M0's gate is where the claim is first tested. If it cannot be met, the gate is renegotiated in this tracker, with the measurement that forced it — never quietly.
R2 The all-reduce may cost more than the reads it hides. One reduction per MoE layer against tens of expert reads: if synchronisation is not small relative to the read time, scaling flattens and the 4–10x thesis is wrong. DC-051 measures the budget at four nodes before any scaling claim is made.
R3 8 GB per node bounds what is resident. Dense backbone, shared expert, KV state and expert cache all have to fit beside macOS. The shard policy and the model choice are constrained by measured resident cost, not by intent (DC-011, DC-032).
R4 Transport reality versus documentation. LAN, SFP/QSFP and a Thunderbolt bridge differ in latency and in what macOS permits; a bridge may need elevated entitlements. Measured 2026-09-15: the farm has two paths — a 1 Gbit Ethernet segment at 0.49–0.64 ms round trip and a mesh VPN at 1.4–1.8 ms — and the node names resolve over the VPN, so a run that binds what a host name resolves to silently takes the slow path. No Thunderbolt bridge is configured on any node, despite two ports each. DC-008 records measurements and the chosen transport; the README's topology claims are corrected if they do not hold. The all-reduce payload is ~4 KB per MoE layer, so this is a latency question before it is a bandwidth one — which makes the 2.5x between the two paths a first-order number, not a footnote.
R5 Licence mismatch. The sister project is Apache-2.0; this repository is MIT. Reusing kernels, the repacker or format code carries Apache-2.0 obligations and a NOTICE. DC-013 decides what is reused and records the obligation before any reuse.
R6 Top-k routing versus 1/N slicing. With k experts selected per token and N nodes, an unlucky slice or routing imbalance leaves nodes idle — and with N greater than experts/k the sharding stops being meaningful. The shard policy is validated on real routing statistics (DC-011, DC-040), and M3's ratio is the honest test.
R7 No reference to be exact against. Bit-matching has no target until a reference implementation is chosen for the M0 model. DC-020/DC-021 pin the model and the reference before the gate is attempted.
R8 Scattered reads on one SSD per node. Sustained throughput on four Mac minis reading expert shards concurrently is not the burst number on one machine. The M1 baseline is recorded on the same hardware class as the cluster (DC-034).
R9 Public-repository hygiene. This wiki and the repository are public: credentials, addresses, access paths and model-access keys must never be written down. Node names are used as labels (node1…node4) and nothing else identifies a machine; the Testbed page records hardware class, measurements and roles only (DC-083 keeps it that way).
R10 Bit-identity across shard counts does not follow from "accumulate in fp32". If each node pre-sums the experts it owns and the partials are then combined, the association order differs from a single-node pass. Measured: 11185 of 20000 random top-8 draws (56%) sum differently — by exactly 1 ULP — under the two orders. The M2 gate would fail on arithmetic alone, with every per-tensor check still green. The reduction must be canonical by construction: nodes exchange their per-expert contributions and every node accumulates in fixed expert-id order, so a single-node pass and an N-node pass perform the same additions in the same sequence. The payload grows to top-k × 4 KB (≤ 40 KB) — affordable, since the transport is latency-bound (Testbed). This is a P0/P3 decision, not an implementation detail.
R11 bf16 router logits flip the top-k set. I3 says discrete decisions must match exactly and keeps routers at "bf16 or higher". Measured over 20,000 random 256-way routers with round-to-nearest-even bf16: 922 (4.61%) changed the top-8 index set. A concrete pair: fp32 logits 5.363673687 and 5.385571480 both round to bf16 5.375, so the ordering that separates them is gone. One flipped index diverges the output completely while per-tensor MSE still looks healthy. (Correction, 2026-09-15: an earlier version of this row quoted a pair at 5.424064/5.413122 — those came from mantissa truncation, which is not what a bf16 cast does. The rate is unchanged in substance; the example was wrong and is replaced.) Resolved as D5: router logits, normalization, comparison and top-k run in fp32, with an explicit ascending-expert-id tie-break applied identically on both sides; bf16 or quantized router weights are fine because the accumulation is fp32. The I3 assertion is a separate test from any numeric tolerance and needs a fixture whose logits sit within 1 ULP of each other, or it cannot fail — test_bf16_rounding_reproduces_the_router_hazard is that fixture.
R12 M0's model is not the conventional one the milestone was designed around. D1 chose Qwen3.5-2B, whose first layer is a chunked Gated DeltaNet: the bit-exactness harness would be first exercised against a chunked delta rule with cumulative decay, a depthwise causal conv and a gated norm — the kernel family the brief assigns to M4/M5 — and, if 4-bit lands in the same step, against a new quantization path at the same time. Two independent error sources in the first proof of the harness is how a harness gets quietly distrusted. The milestone is staged (M0a harness + conventional model, M0b bf16 on the real architecture, M0c 4-bit) so exactly one variable moves at a time; the work is not wasted, because M1's Qwen3.6-35B-A3B is 30-of-40 layers of the same Gated DeltaNet family, and the sister project already runs that family single-node as a comparison point.
R13 PyTorch's matmul order is not reproducible by any other implementation. Measured on a real layer's shapes (8×1024 by 512×1024, fp32): an explicitly ordered accumulation differs from torch's in 3544 of 4096 outputs (mean 17 ULP, max 22587 ULP), while both sit exactly 3.148e-07 from an fp64 computation of the same product. The difference is summation order, not accuracy, and the order belongs to torch's BLAS kernels rather than to the model. A gate defined as "bit-match torch" would fail forever, for a reason that has nothing to do with the engine. Resolved in D3 by splitting the references: tools/ordered_reference.py states the order of every sum and is the bit-exactness target, while the torch trace is the semantic oracle — exact discrete decisions (I3) and per-tensor closeness. The engine's kernels must accumulate in the contract's order (no split-K, no FMA, no reassociation), and I2 compares our implementation to itself. Confirmed on the real model: the ordered forward's logits argmax agrees with torch at every position while the last bits differ from the first matmul onward.

Open questions

Decisions that are not yet made. Each one is answered by an ADR or a task above, and the answer is recorded here even when it is "not yet decided".

ID Question Answered by
Q1 Which dense model is M0's target, and which implementation produces its golden traces? DC-020, DC-021
Q2 Is the IR a shared component with the sister runtime, or a new one built beside it? DC-010
Q3 How does the 1-node M1 baseline run the sharded model for a bit-identical comparison — one shard holding every expert, or a separate unsharded path? DC-006, DC-035
Q4 What exactly stays resident on each 8 GB node, and what is streamed? DC-011, DC-032
Q5 Which transport is mandatory for M2/M3, and does the Thunderbolt bridge need elevated privileges? (Measured: 1 Gbit Ethernet 0.49–0.64 ms vs mesh VPN 1.4–1.8 ms; no bridge is configured — see R4 and DC-017.) DC-008
Q6 Does "unlimited nodes" mean N is bounded by experts-per-layer divided by k, or is expert duplication allowed? DC-011, DC-050
Q7 Where do golden traces live — in the repository, or in external storage? Trace size decides it. DC-007
Q8 Is Core ML / Neural Engine prefill part of a node's engine, and if so where is the sidecar produced and verified? (The conversion toolchain is installed — DC-016 — but installing it decides nothing.) DC-033, DC-061
Q9 Why 2,218 expert requests for a five-token prefill? That is 1,109 fetches, or 221.8 per token, against 8 experts × 40 layers = 320 selections per token. De-duplicating selections within a layer would land below 320, but that is an inference: settle it by counting distinct experts per layer in the mixture and comparing DC-034

Later / parked

Ideas that are not scheduled and not promised. They are written down so they are not lost, and they are not tasks until they get a DC-nnn id.

  • A docs/ tree in the repository for runbooks and audit registers, mirroring the sister project's split between wiki (user/engineering documentation) and docs/ (runbooks).
  • Package the cluster launcher as a single script that reports per-node readiness before a run, in the spirit of the sister project's launcher.
  • A public benchmark page once there is a number worth publishing — with the prompt set and the hardware named, or not published at all.
  • Report the sister project's SECURITY.md defect upstream: its private vulnerability reporting link points at github.com/drumih/turbo-fieldfare, not at TinyTitan. Found on 2026-09-15 while writing this repository's own policy; it belongs in the sibling tracker or an upstream issue (DC-085).
TinyTitan Datacenter

Project

Design

Development

Sister project: TinyTitan

Clone this wiki locally