Repository navigation
Project Tracker
| Phase | P2 — M1 (P0 in progress; P1's M0 gate passed) |
| M0 |
Gate passed — Qwen/Qwen3.5-2B revision 15852e8c…, three frozen prompts, 40,683,520 bytes of trace data identical to the contract and every discrete decision matching the reference (docs/m0-gate.md) |
| M1 |
Gate passed — and now verified against a contract reading the same weights (2026-09-17). trace_diff between the engine and a contract run on the install itself reports IDENTICAL — 83 tensors, 0 differing elements, 40 discrete decisions, matching digests b0d382dbabf36df0…. The claim is restated in tools/milestones.json to that falsifiable form, because the earlier comparison had two inputs: the engine read an install and the contract read a bf16 checkpoint, and D55 showed the whole divergence was that difference — the DeltaNet's three projections and attn.q/k/v/o are int4 in the install, which moves layer 0's attention output by 0.0156 max / 24.5% median relative. Against the checkpoint contract the pair differs by 40 discrete decisions and 1 float tensor, declared and tracked in DC-112. Engine measurements that stand: 0.108 tok/s cached and 0.0374 uncached generation with identical tokens, a hit rate of 0.0000 over 2,240 requests (structural), and 348.6 MB peak memory |
| M2 |
Gate passed, and re-established on 2026-09-17 against the current binaries. A 256-expert plan over two contiguous halves, one node on a peer machine over TCP and one here, each reading its own install half (2.05 GB apiece), 40 reductions per node over 797 and 803 terms: trace_diff reports IDENTICAL — 83 tensors, 0 differing elements, 40 discrete decisions, matching digests b0d382dbabf36df0… for both nodes against the single-node reference. This matters because the engine has changed since the original run — the GPU unpack became the default and every contract matmul went through a chooser — and both are asserted bit-identical, which is exactly the sort of assertion that wants re-checking rather than trusting (D66). |
| Open tasks | 35 across the phases below; closed work and the dated record are on News |
| Language standard |
Swift 6.4 on Xcode 27 with swift-tools-version:6.4 — the only supported toolchain, enforced by tools/check_toolchain.py; Python 3.14
|
| Tests |
226 Swift, 427 Python, both run on every push — and nothing is skipped: the Metal kernel tests run on the node's GPU since D34
|
| Farm | Four identical Mac mini M2, 8 GB each, macOS 27.0, Xcode 27, Python 3.14.7 |
| Development link | 1 Gbit Ethernet: 0.49–0.64 ms round trip measured; the mesh VPN path measures 1.4–1.8 ms |
| Phase | Milestone | Objective | Gate | Status |
|---|---|---|---|---|
| P0 | Foundation | Decisions, harnesses and conventions | ADRs, harness spec and conventions written | In progress |
| P1 | M0 | Staged: harness (M0a), Qwen3.5-2B bf16 (M0b), 4-bit (M0c) | Bit-matches reference golden traces at each stage | Done |
| P2 | M1 | Qwen3.6-35B-A3B, single node, 4-bit, SSD-streamed | Correct output + recorded tok/s baseline | Done |
| P3 | M2 | 2 nodes, expert-parallel | Bit-identical to M1 | Done |
| P4 | M3 | 4 nodes | ≥3x the M1 tok/s | In progress |
| P5 | M4 | DeepSeek-V4.1-Flash | Matches reference at 128K context | Planned |
| P6 | M5 | Qwen3.8-Flash-Next | Same protocol as M4 | Planned |
-
Task IDs (
DC-nnn) are stable and are never reused or renumbered. A cancelled task keeps its ID and is markedDroppedwith the reason. -
Status is one of
Planned·Next·In progress·Blocked·Done·Dropped.Nextmeans it is the immediate queue, not that it has started. - "Done when" is the evidence, not the activity. A task is done when the named artifact or measurement exists, and the row says where it is.
- A gate is closed only by measurement. When a status here disagrees with a gate on the Roadmap, the measurement wins and this page is corrected.
- Documentation sync rule (DC-082). A commit that changes behaviour, the plan, or a decision updates the README, the affected wiki page and this tracker in the same push. A tracker row that has drifted from the repository is a defect.
Decisions that are expensive to reverse. Each ADR records the alternatives considered and why the others were rejected.
| ID | Task | Depends on | Done when | Status |
|---|---|---|---|---|
| DC-116 | Establish whether dense models belong to this design, and record the answer | The target was an overview across Qwen3.5 4B, 9B and Qwen3.6 35B-A3B, so a 4B install was built from the real checkpoint (426 tensors, 4,204,789,760 weights, 3.34 GB at 6.36 bits/weight) and measured on one node. Dense is out, and the measurement says why: the shard plan divides experts, so a dense model has no plan at all and cannot be distributed; and its whole payload is re-read for every token — 18,638,208,000 bytes in four steps, with 0 hits, because the default 1 GiB SHARD_DENSE_CACHE_MB cannot hold a 3.34 GB payload. Decode came out 5.870 s/step = 0.170 tok/s, i.e. slower than the 35 B MoE on the same node (0.230 tok/s), which activates only ~3 B of its 35 B. The 9 B is worse still: 5.5 GB of payload against ~4.5 GB of usable RAM, and still undividable. The table carries the 4 B row with its prefill unmeasured, and the README says dense is outside the design |
The answer is recorded with its number, D96 carries the reasoning, and the README's table shows the MoE rows beside the dense one rather than leaving a reader to wonder why 4 B and 9 B are absent |
Done |
| DC-115 | Cut the first release (v1.0.0), with an identity a build can enforce and an archive a user can verify Done and verified 2026-09-17: v1.0.0 is published with the archive and its checksum, the notes quote digest b6ebfd1b861a1ac4… and the .sha256 beside the archive matches it, shasum -c passes on the downloaded bytes, and all three shipped binaries answer 1.0.0 and report arm64. One blemish, on the record rather than rewritten: the archive also carries two test bundles, because the first version carried every bundle the build directory held instead of the set the plan says the executables can reach — harmless, 2 MB of fixtures in bin/, and wrong. The selection is now plan-driven (declared_bundles, with tests) and reports (none declared) for this package; the published assets are left alone, because rewriting a tag someone may already have downloaded is worse than a blemish that is written down. |
RELEASE.md §1.3 wants the version single-sourced and enforced, and §1.6 wants the archive to carry what a Swift binary needs. There was no VERSION, no --version, no release script and no changelog, so the identity had to exist before a release could. tools/release.py runs the gates, does a clean scratch build with the log scanned for warnings, packages the three executables with the licence, the notices and a README-binaries.txt, asserts lipo -archs on the binaries extracted from the archive, and refuses to publish notes that do not quote the digest it just computed (D95) |
A published release whose assets are the archive and its checksum, with --version answering 1.0.0 from the binaries inside it, and tools/version.py --check refusing a drifted mirror |
In progress |
| DC-114 | Overlap the exchange wait with the next layer's decode Not needed for the objective 2026-09-17 (D94): the cluster reached 1.74x without overlapping the wait, so this row's done-when is now about the roadmap's ≥3x rather than 1.5x. |
The cluster is 4.021 s/step and needs 3.658 for 1.5x, and its largest single phase is load at 1.9 s/step — 47.9% — during which the GPU idles and the CPU dequantises constants (D88). A node waiting for a peer's contributions is idle too, and the work it will need next is known: decoding the following layer's weights during the wait turns one into the other. It must be one layer of decoded weights held transiently, not a whole-model cache: D89 measured that holding 16 layers loses to memory on an 8 GB node, and the difference is that a one-layer prefetch holds bytes the next layer needs anyway. |
A cluster step at or below 3.658 s/step with bit-identity intact, or a measurement showing the wait is too small to hide anything in | Open |
| DC-113 | GPU decode: a quantized GEMV that dequantises in-kernel Decode is M=1, so the right GPU kernel is a GEMV that reads int4 weights and dequantises inside it. D88 measured load at 30.5% of a cached step — the same ~1 G parameters of constants dequantised on the CPU on every token — and D89 showed caching them loses to memory on an 8 GB node, so the phase has to stop existing rather than be paid for twice. It can stay bit-exact (D63: one thread per output, k accumulated in order — tiling changes which thread works, not the order of the sum), which an ANE path cannot (D91). Any device selection must verify the device it got and refuse a silent CPU fallback: the sister project's measured trap is Core ML exiting 0 while running on the CPU at ~38x the GPU cost (D91). Superseded on the path to this number 2026-09-17 (D94): the CPU dequantiser's row loop now runs across the machine's cores, which took the same phase down 2.4x (1.33 → 0.56 s/step) with no device, no protocol, no bit-exactness argument to make and none of the silent-CPU-fallback risk D91 found. The GEMV remains a candidate for absolute speed, not for this objective. |
— | The decode step's matmuls run on the GPU from quantized weights, bit-identically, and the step is measurably faster | Open |
| DC-008 |
ADR — transport: Thunderbolt bridge vs LAN vs SFP/QSFP, measured latency and bandwidth per hop. The implementation half is done (D22): a listener that binds, listens and accepts, a connector with an explicit deadline, and a contribution exchange over TCP that is bit-identical to the single-node forward. What remains is the half the ADR is named for — the per-hop measurement — which needs the testbed, plus IPv6 if the farm needs it |
DC-009 | The tabled latency and bandwidth come from measurements on the real links, not from the paper | In progress |
DC-004 status: CI carries the repository's own gates on every push — the Markdown link
gate (tools/check_markdown_links.py: local links and #anchors, offline), the table gate
(tools/check_markdown_tables.py, over the repository and this wiki) and the whole tools
suite. The Swift workflow (.github/workflows/swift.yml) fails on a runner that cannot
meet the project toolchain instead of warning and skipping (DC-103): GitHub's macos-26
image carries Xcode 26.0–26.5, so that job is red by design until an image ships Xcode 27,
and the Swift gate runs locally until then. The tools runner has no torch, so the capture
tests skip there.
Closed in this phase: DC-001 Bootstrap the wiki: Home, Roadmap, Project Tracker, · DC-002 README correction pass: the Shard codename in the · DC-003 Repository scaffolding: .gitignore, AGENTS.md, · DC-006 ADR · DC-007 Golden-trace harness specification: trace schema, · DC-010 ADR · DC-012 Repository layout and SwiftPM target plan, consistent · DC-014 Language standard recorded and applied: Swift 6.4, · DC-015 The Swift language-feature register: which upcoming · DC-016 Core ML tooling installed and verified: coremltools · DC-017 Farm inventoried from this checkout: every node's · DC-019 Interconnect, storage and memory baselines measured, · DC-037 M1's numeric contract: the mixture of experts · DC-038 Expand the key and query heads up to the value head. Evidence on News.
The harness is proven before the model, and the model is proven in the order that keeps one variable at a time (D1):
Closed in this phase: DC-020 Choose and pin the M0 model: Qwen3.5-2B (Apache-2.0) · DC-021 Reference runner that emits golden traces for the · DC-022 IR importer for the M0 models (name-to-role mapping · DC-023 Minimal single-node forward pass, driven by the IR, · DC-024 Trace-capture harness (per-layer activations, logits, · DC-025 Diff harness with bit-exact comparison and · DC-026 Gate M0: on the pinned model, zero differing bytes · DC-027 Reference contracts for both M0 models: the official · DC-028 M0a: capture and diff harness, proven able to · DC-029 M0b: Qwen3.5-2B bf16 bit-exact, layer by layer,. Evidence on News.
| ID | Task | Depends on | Done when | Status |
|---|
Closed in this phase: DC-030 Importer for the Qwen3.6-35B-A3B · DC-031 Quantization/transcode · DC-032 SSD expert streaming with a bounded per-node expert. Evidence on News.
Sweep result (2026-09-16). capital's five tokens at bank sizes 2, 8 and 16 produced three
byte-identical traces — trace_diff reports 83 tensors, 0 differing elements and 40 discrete
decisions — with identical expert counters at every size: 6,977,224,704 bytes from SSD (1,395 MB per
token), 2,218 requests, 0 hits, 13.95 GB decoded in memory, and 39.3–39.9 s of prefill (7.90 s per
token, 0.127 prefill tok/s — a prefill rate, not the gate's generation baseline). The digest moved
from the pre-D11 b8c976c5… to b0d382db… exactly as D11 predicted. Full numbers and the one
unexplained figure (2,218 requests against 320 selections per token) are in docs/m1-gate.md.
| DC-107 |The kernels are done and they are the wrong lever — the engine is I/O-bound (D64). The tiled kernel is bit-exact (288-shape grid, head shape, cache reuse; the real trace b0d382db… either way, trace_diff IDENTICAL) and still slower than the CPU: attn.core 4.05 → 5.6-5.9 s, mix.gateup 1.18 → 2.1 s. Coalescing the loads changed nothing, so the cost is the path around the kernel — an encoder, two copies and a synchronous wait per call, ~4 ms across ~210 calls in mix.gateup — not bandwidth. And the profile says where a forward goes: mix.read + load + head are 12.73 s of 18.75 s, 68%, at the ~1 GB/s sequential floor. So the lever for M3's ≥3× is sharding the reads (M2/M3 already do), not arithmetic. The kernels stay opt-in (SHARD_GPU_MATMUL=1) as a tested asset. Remaining targets here: the reads, and attn.core's 4 s of compute, both bounded by the same 68%. The GPU matmul is bit-exact and SLOWER, so it is opt-in (D63) — and the head's 1.74 → 1.28 s was drift. Routing every contract matmul (38 call sites) through the chooser leaves correctness untouched — the trace is b0d382db… on the GPU and on the CPU, trace_diff IDENTICAL — but with the conditions alternated rather than run in sequence, every phase that uses it is slower: attn.core 4.05 → 5.55 s, mix.gateup 1.17 → 1.62 s, head 1.74 → 1.86 s, while mix.read (no matmul) is unchanged, so it is the kernel and not the machine. Cause: one thread per output walks w rows k*4 bytes apart, so a warp touches 32 cache lines for 4 useful bytes each; the CPU's path walks k contiguously. Fix: a threadgroup-tiled kernel, which cannot change a result because tiling changes which thread accumulates, not the order. SHARD_GPU_MATMUL=1 until then. The head still reads 1.02 GB of its own weights, which is its floor. The head is done and it was half compute, half I/O (round 45). The GPU matmul — one thread per output, ascending k, acc + metal::fma(x, w, 0), the spelling D61 measured — is bit-identical to Ops.orderedMatmul over 288 shapes (including output counts that are not multiples of four, where the CPU vector takes its tail), on the head's real shape and across cache reuse; the real trace is b0d382db… with it on and off, trace_diff IDENTICAL. Measured by phase: head 1.74 s → 1.28 s, while the totals said nothing (16.3 against 16.4 s, the wrong way round, because the reads vary more than the saving). The rest of the head is 1.02 GB of head weights at ~0.8 GB/s — the same sequential floor as the expert reads — so it is documented as at its limit. attn.core 4.05 s did not move and is the opposite case: small activations, arithmetic-dominated, and the right place for this kernel. The head kernel's arithmetic is now settled (D61). Before writing it, the question D10 left open was measured: .fast, .relaxed and .safe all contract a * b + acc into one fma (.safe only keeps a product in its own statement apart), so no mode is enough — a GEMM must write accumulator + metal::fma(x, w, 0.0f), which reproduces the contract's separate rounding under all three modes and lets it share .relaxed with the unpack. The probe uses a one-ULP discriminator for that reason: the same probe over a 257-term dot passed in every combination, so a long-dot test would have declared the GPU exact and been wrong. Re-profiled (round 43), and the order changed: mix.read 7.23 s of a 19 s trace — the install read 3.24 s plus the unpack 4.53 s — then attn.core 4.05 s, load 2.22 s, head 1.76 s, mix.gateup 1.22 s, mix.down 0.53 s. The unpack is no longer a target: the GPU path is now the default, verified bit-identical three ways, and takes about two seconds off the trace (D59, D60) — so the earlier note that it was "at its measured limit" was wrong and is corrected here. The read genuinely is at its limit: 3.24 s for 3.06 GB is ~0.95 GB/s, the sequential floor, and no kernel changes that. The next kernel should be the head, a plain GEMM (1.76 s) and the easiest of the remaining ones to make bit-exact, before the DeltaNet's chunked rule inside attn.core. The unpack target is met (round 42). The GPU unpack is now faster than the scalar one with the same digest: 16.3 s against 38.2 s before the buffer cache and 18.2 s on the CPU, with trace_diff against the scalar run reporting IDENTICAL — 83 tensors, 0 differing elements, 40 discrete decisions. The fix was the plumbing the diagnosis named: a one-slot buffer cache held under a lock across the dispatch and its completion, and one copy per section instead of four. Recorded as wall-clock observations of one trace each, not a benchmark; the default stays opt-in until the timing phase, when the frozen prompt set can be the evidence (D59). The remaining targets here are attn.core and the reads. Round 41 measured the unpack's GPU path: same digest, slower — 18.2 s scalar against 38.2 s GPU for the five-token trace. The kernel is not the cost: the row path concatenates codes + scales + zeros and unpack copies it again, and a fresh MTLBuffer is allocated per expert fetch (2,218 of them). So the remaining work here is the plumbing — a reusable buffer, or a dispatch over a payload already contiguous — not the shader. The next performance targets: attn.core 3.86 s, the expert unpack 3.45 s, the reads 2.99 s. After D15 and D16 removed 24 s of hashing and duplicate reads, these are what a 14.97 s five-token forward is made of — none of them I/O, and the reads and unpack are at their measured limits (~1 GB/s and ~1,012 M values/s). The attention core is the largest single compute phase and the first target | DC-105 | Each is either faster with the same digest, or documented as at its limit with the measurement that says so | Planned |
| DC-112 |Decided (round 40): the DeltaNet's and the attention's projections should return to bf16. The logits settle it: against a bf16 checkpoint contract the argmax is unchanged on all five positions, but the smallest margin is 0.27, top-8 overlap is only 5–8 of 8, and the mean logit difference is 0.30. That is the same situation the policy already treats as decisive for the router and the gates — they decide the discrete outcomes — and a narrow pass is not a reason to spend it. The routed expert stacks (~19 GB of the 21.7 GB install) stay int4, which is where the size win lives. Execution is blocked here: the rebuild needs ~25 GB free and this node has 9 GB, and it is a GB-scale job that must not run on this machine. The gate conflict is settled (round 39); what remains is a quality-and-size choice, not a correctness one. The milestone's claim is restated to the form that can be falsified and proved: the engine's trace is byte-identical to a contract reading the same install (IDENTICAL — 83 tensors, 0 elements, 40 discrete decisions, matching digests), verified 2026-09-17. The design decision left open is whether the DeltaNet's and the attention's projections stay int4 — about 3 GB of a 21.7 GB install, with the routed expert stacks' ~19 GB unchanged either way — or return to bf16 to narrow the gap to a checkpoint contract, now measured at 0.0156 max / 24.5% median relative at layer 0. A rebuild to test that needs ~25 GB free and must not run on this node. Round 39, the way to settle it without a rebuild: give both sides the same weights. D55 showed the whole divergence was weights, not arithmetic, so the gate's comparison had two inputs and could not falsify anything. tools/install_source.py feeds the contract from the install through the same dequantiser the Swift reader mirrors, so a disagreement is a disagreement in arithmetic. Findings: an install flattens trailing dimensions, so one expert is 1,024 consecutive install rows of the 262,144-row expert.stack_gate_up — D32's "one expert is one row" is true of the checkpoint, not the install; Install.row_range makes expert 255 cost one expert; ten new tests (359 total). The full-model run is in flight — 1,600 expert fetches through Python, 878 MB and ~19.5 min CPU at last measurement, both guards live. Not claimed until measured. Settle the conflict between M1's gate and the quantisation policy (raised by D55). Either keep the DeltaNet's and the attention's projections at bf16 — roughly +3 GB on a 21.7 GB install, since the routed expert stacks are about 19 GB of it and stay int4 — or restate M1's claim with the measured cost of int4 as its evidence: 0.0156 max / 24.5% median relative at layer 0. Rule 3 forbids the quiet version of either. Blocked on resources here: the rebuild needs about 25 GB free and this node has 9.5 GB, and it is a GB-scale job that must not run on this machine at all | DC-111 | The decision is recorded, and either the policy or the claim changes with the measurement beside it | Open |
| ID | Task | Depends on | Done when | Status |
|---|
| ID | Task | Depends on | Done when | Status |
|---|
| DC-051 | Per-token synchronisation budget: all-reduce latency per MoE layer against expert-read time Measured 2026-09-17 (D84): the budget is 0.893 s/step of exchange — 15.0% of a 5.975 s step — moving 4.49 MB at 5.0 MB/s effective, i.e. 1.63 ms per exchange term over 547 terms/step and 90 reduces/step, which is a per-message latency rather than a saturated link. That accounts for the exchange; it does not account for the ~3.8 s/step that is neither payload read across the plan nor exchange, because the node reads 2.65x less than the baseline and is slower. Remaining: a per-phase breakdown of that 3.8 s, which is the instrument this row and DC-107 still need. Step-level budget measured 2026-09-17 (D88), on the cached decode the gate measures: mix.read 1.80 s/step of device reads at 184 MB/s (0.33 GB/step), load 1.66 s/step of CPU dequantisation of constants, head 1.05 s/step of matmul on cached weights, attn.core 0.52 s/step. The exchange (D84) is 0.89 s of the cluster's step. The plan therefore addresses 33% of a step, and the target is not reachable from it. The cluster's own budget measured 2026-09-17 (D90), per node: the exchange is 4.250 / 4.136 / 0.672 / 3.739 s/step on nodes 0-3 — 71.5%, 69.5%, 11.2%, 63.1% of the step — while the plan is demonstrably working (reads 6.281 → 2.31-2.40 GB per node, mix.read 1.80 → ~0.49 s/step). A 6× spread between nodes doing identical work is the signature of head-of-line blocking: allReduce sends to every peer and then receives sequentially in peer order, blocking, so one slow peer delays everyone behind it, forty times per token (~105 ms per layer). The fix is the budget's other half: receive concurrently and skip peers that own none of the chosen experts. And the counter does not reconcile with the phases (2026-09-17, D93): the ledger's exchange_seconds is 1.4-1.9 s/step on every node while the ff phase that contains it is 0.575 s, even though ff is marked after mixtureOutput returns and the phases sum to the step (3.9 of 4.02 s). Both cannot be right, and D90/D92 reasoned from that counter, so it is recorded rather than used until it is resolved. | DC-045 | The measured budget explains the achieved scaling | Planned | In progress
| DC-052 | Per-node cache and prefetch tuning at four nodes Target identified by measurement 2026-09-17 (D88): the cached step's load phase is 30.5% of a 5.449 s step — 1.66 s/step — and it is not I/O. The dense payload is read once for the whole run (1.0437 GB held, 4888 cache hits), so that time is the same ~1 G parameters of constants being dequantised and released on every token, 160 times per generation, on the CPU while the GPU idles. mix.read (33.0%, 1.80 s/step, 184 MB/s) is the only device-bound phase and the only one a plan divides. So this row is the speed lever: decode once and hold, with the budget measured and recorded per node rather than guessed. Built and measured 2026-09-17 (D89): the decoded-layer cache holds layers across the steps of a generation — not an LRU, which a sequential sweep gives a zero hit rate by construction — counts the bytes it holds, and records budget/held/hits/misses per node in metrics.json. It is bit-identical (asserted tensor by tensor) and it works: about 0.035 s per layer per step, 1.4 s across all 40. On this 8 GB node it still loses — at 1 GB and 2 GB the step rose to 5.46 and 5.60 s from 5.33 s, with head +0.20 s/step and attn.core +0.27 on a machine already swapping — so the default budget is 0, and SHARD_LAYER_CACHE_MB is the knob for a node with headroom. Prefetch depth remains open. | DC-050 | Cache budget and prefetch depth are recorded per node | Planned |
| DC-053 |Round 59: the INSTALL form of M1's gate also passes, all five frozen prompts. DC-114 was fixed — the install dequantiser is vectorised, bit-identically, 0.253 s to 0.013 s for one real expert — and the gate that could not finish a five-token prompt in twenty-five minutes now completes the whole set: capital 17.40 s, arithmetic 54.51 s, code 67.63 s, repeat 86.90 s, long 98.45 s, every one IDENTICAL — 83 tensors, 0 elements, 40 discrete decisions — GATE PASSED, digests b0d382dbabf36df0… on both sides, peaks 0.98-1.41 GB. That is M1's restated claim (D56) checked by its own instrument, alongside the checkpoint form checked in round 56 (D73). The install declaration moved 1.5 -> 2.0 GB because the measured peak is 1.41 GB and a 6% margin is a coincidence rather than a guard. Round 56: M1's gate PASSED on the real checkpoint, on this node, all five frozen prompts. The checkpoint path was still going through safe_open — mapping 67 GB, the hazard the node's rules name — because the gate passed only one of the two flags docs/m1-gate.md records as the survivable pair; it now passes --uncached on the checkpoint path too. Measured first: the contract on the checkpoint is 66.20 s and 0.397 GB peak, against 4.16 GB recorded before the flags. Then the gate: capital 34.94 s / arithmetic 89.65 s / code 115.92 s / repeat 128.13 s / long 146.26 s, all IDENTICAL — 83 tensors, 0 elements, 40 discrete decisions — GATE PASSED, peaks 3.60-3.79 GB, and the first prompt's digest b8c976c5e7ba8816… is the one docs/m1-gate.md recorded on 2026-09-16. The declaration for this path is 4.2 GB, restored from a 1.5 GB I set from the contract's measurement until the gate's own 3.79 GB peak corrected me: a checkpoint run's peak is the engine's trace, not the contract's (D73). What remains here is the install path (DC-114) and the ≥3× measurement in a quiet window. Round 51: the four-node mesh is re-verified and the functional half is runnable on its own. With the farm quiet (node1 2.5, node2 1.4, node3 1.0) run_m3_gate.py --functional-only ran the full mesh across four machines against the single-node run: every node IDENTICAL, tokens and digests, over a plan staged to three peers that already held their own installs. The new flag expresses the operator's own split — it checks the cluster and stops, computing, printing and recording no speedup, and nulling the timing fields rather than omitting them so the output cannot be mistaken for a measurement; it also skips the quiet-farm precondition and says so, because a functional run makes no throughput claim. What remains here is the ≥3× measurement in a quiet window, which is deliberately not run yet. Round 48: the tokens are verified and now gate-checked; the gate itself still needs a bigger node. The cached and full-sequence paths generate the identical eight tokens (11751,11,264,3177,34756,364,1141,8807) with identical top-2 margins at every step — including 0.1339 and 0.0445, which are narrow enough that a numerical difference could move them — and run_m1_gate.py now runs both and requires equality (I3, no tolerance), with an empty parse failing rather than passing. The gate's own fixture test exercises the new path. What remains for this task is unchanged and is the throughput measurement in a quiet window; Corrected in round 52: run_m1_gate.py now discovers whether it was handed a checkpoint or an install and runs both halves against it, and it passes --stream-experts to the contract — without which the contract half could not read an install's expert stacks at all (D69). The install form needs no 70 GB mapping, so M1's own gate is runnable on this node; what D65 said about the checkpoint form still holds. Gate M3: ≥3x the frozen M1 tok/s on the same prompt set. The functional half is demonstrated (D27: four machines in a full mesh, one bit-identical trace) and the instrument now exists (D28: datacenter-generate shards, and a two-machine cached generation produced the reference's tokens and digest). What the gate needs is the measurement, on a quiet farm, since the nodes are shared with other work First real measurement 2026-09-17, on a busy farm with the operator's authorisation (--allow-busy-farm, so the gate recorded it as observation_only with the loads and did not assert the threshold): bit-identity holds on all four nodes, and the ratio is 0.93x — baseline 5.662 s/step against the slowest node's 6.086 s/step. The per-node metrics show why: a node reads 2.65x less payload per step (0.592 GB against 1.570 GB, because the dense 0.261 GB is replicated and the experts really are 3.95x fewer) and is nevertheless slower, and the exchange is 0.893 s/step (15.0%) moving 4.49 MB at 5.0 MB/s effective — 1.63 ms per term, latency-bound, not bandwidth-bound. So the >=3x target is not what an expert plan delivers on this design: the step is not expert-read-bound. The threshold remains unasserted and this is not a gate result — four steps, a busy farm, and the quiet-window measurement is still owed (D84). And measured 2026-09-17 (D94): the cluster runs at 1.74x and 1.70x over a single node with eight decode threads, and at 1.66x with one — four alternated runs, every one bit-identical on all four nodes, loads 1.6-5.5. The operator's objective (≥1.5x) is therefore met on a busy farm, and load hurts the ratio (four nodes are exposed to a spike and the baseline is one), so these are lower bounds. The roadmap's own ≥3x on a quiet farm is still not asserted, and the gate remains the certification instrument. | DC-045 | On a quiet farm, the same prompt set runs at ≥3x the M1 rate with bit-identity intact | In progress |
| ID | Task | Depends on | Done when | Status |
|---|---|---|---|---|
| DC-060 | DeepSeek importer with transcoding of native FP4 experts and FP8 dense weights | DC-045 | No requantization step exists in the path | Planned |
| DC-061 | Compressed / sparse hybrid attention and constant-size recurrent state | DC-060 | The model runs at the reference context length | Planned |
| DC-062 | MTP speculative decoding parity | DC-060 | Draft and verify match the reference protocol | Planned |
| DC-063 | Sparse-attention block-selection exactness check | DC-061 | Selections match the reference exactly, not approximately | Planned |
| DC-064 | Gate M4: matches the reference at 128K context | DC-063 | The exactness protocol passes at 128K | Planned |
| ID | Task | Depends on | Done when | Status |
|---|---|---|---|---|
| DC-070 | Importer and reference match for Qwen3.8-Flash-Next | DC-064 | Same exactness protocol as M4 | Planned |
| DC-071 | Sparse n-gram lookup table handling and per-node sharing | DC-070 | Table lookups are correct and bounded in memory per node | Planned |
| DC-072 | Expert-parallel scaling on the 125B-A6B shape | DC-070 | Gate M3's ratio is reproduced or explained on this shape | Planned |
| DC-073 | Gate M5: matches the reference, same protocol as M4 | DC-072 | Recorded, with the measurement | Planned |
| ID | Task | Depends on | Done when | Status |
|---|---|---|---|---|
| DC-085 |
Re-audited 2026-09-17 against the same snapshot, and the position holds with evidence. Its int4 is unsigned with a bias, and its own reference says so in the first line of the function that consumes the format: its own qwen35_reference.py's dequantize() — "bits-wide unsigned lanes packed low-first inside each uint32, one BF16 scale and bias per group" — computing grouped * scales + biases, with dequantizeInt4Affine on the Swift side. Ours is signed codes with an int8 zero point (_signed() and (code - zero) * scale in tools/install_reader.py), and its container family is GTurbo*V1 (GTurboFormatV1, GTurboExpertV1, GTurboLayerV1, GTurboManifest*V1) against our install.json plus data.bin. The same bytes mean different numbers under the two rules, so a file from one is not readable by the other — arithmetic rather than opinion. No defect to report upstream, because its reader and converter agree with each other. The checkout carries no git history, so "track its releases" cannot be done from it; the snapshot is a source drop dated before the original audit, and the position is recorded with its evidence in THIRD_PARTY_NOTICES.md (D80). Audited 2026-09-16: the sister project's int4 is unsigned-with-bias and its container family is GTurbo*V1, so there is no defect to report and nothing of ours was copied (D39). Synchronisation with the sister project: track its releases, report cluster-relevant defects upstream |
— | A defect found here has a linked upstream issue where it belongs there | Planned |
Closed in this phase: DC-082 Documentation sync rule: README, wiki and this · DC-086 Uncached reads for model payloads: open the payload · DC-088 Per-slab digests, so a reader can verify the expert · DC-089 D11 option 1: flush denormal scales to zero in both · DC-093 WITHDRAWN · DC-094 The spec/config correspondence is a tool with tests · DC-095 The wiki's tables have never been gated in CI, and · DC-096 Audited the two largest unprobed engine surfaces and · DC-097 The brief's 16 KB alignment is unimplemented, and · DC-098 I6's two gaps are closed, all three items verified. · DC-099 Audited I4 and L2 and found both · DC-100 The M1 sweep is a script that refuses to start, not a · DC-101 The brief's own · DC-102 The public status claim is corrected in all five · DC-103 The toolchain is enforced, not documented: Xcode 27 · DC-104 Corrected the M0 status claim in six documents: M0 is. Evidence on News.
| ID | Risk | Handling |
|---|---|---|
| R1 | Bit-identity may not be attainable on Metal. Nondeterministic reduction order or driver scheduling can make two runs differ. | The deterministic reduction contract (DC-005) is written before M0, and M0's gate is where the claim is first tested. If it cannot be met, the gate is renegotiated in this tracker, with the measurement that forced it — never quietly. |
| R2 | The all-reduce may cost more than the reads it hides. One reduction per MoE layer against tens of expert reads: if synchronisation is not small relative to the read time, scaling flattens and the 4–10x thesis is wrong. | DC-051 measures the budget at four nodes before any scaling claim is made. |
| R3 | 8 GB per node bounds what is resident. Dense backbone, shared expert, KV state and expert cache all have to fit beside macOS. | The shard policy and the model choice are constrained by measured resident cost, not by intent (DC-011, DC-032). |
| R4 | Transport reality versus documentation. LAN, SFP/QSFP and a Thunderbolt bridge differ in latency and in what macOS permits; a bridge may need elevated entitlements. Measured 2026-09-15: the farm has two paths — a 1 Gbit Ethernet segment at 0.49–0.64 ms round trip and a mesh VPN at 1.4–1.8 ms — and the node names resolve over the VPN, so a run that binds what a host name resolves to silently takes the slow path. No Thunderbolt bridge is configured on any node, despite two ports each. | DC-008 records measurements and the chosen transport; the README's topology claims are corrected if they do not hold. The all-reduce payload is ~4 KB per MoE layer, so this is a latency question before it is a bandwidth one — which makes the 2.5x between the two paths a first-order number, not a footnote. |
| R5 | Licence mismatch. The sister project is Apache-2.0; this repository is MIT. Reusing kernels, the repacker or format code carries Apache-2.0 obligations and a NOTICE. | DC-013 decides what is reused and records the obligation before any reuse. |
| R6 | Top-k routing versus 1/N slicing. With k experts selected per token and N nodes, an unlucky slice or routing imbalance leaves nodes idle — and with N greater than experts/k the sharding stops being meaningful. | The shard policy is validated on real routing statistics (DC-011, DC-040), and M3's ratio is the honest test. |
| R7 | No reference to be exact against. Bit-matching has no target until a reference implementation is chosen for the M0 model. | DC-020/DC-021 pin the model and the reference before the gate is attempted. |
| R8 | Scattered reads on one SSD per node. Sustained throughput on four Mac minis reading expert shards concurrently is not the burst number on one machine. | The M1 baseline is recorded on the same hardware class as the cluster (DC-034). |
| R9 | Public-repository hygiene. This wiki and the repository are public: credentials, addresses, access paths and model-access keys must never be written down. | Node names are used as labels (node1…node4) and nothing else identifies a machine; the Testbed page records hardware class, measurements and roles only (DC-083 keeps it that way). |
| R10 | Bit-identity across shard counts does not follow from "accumulate in fp32". If each node pre-sums the experts it owns and the partials are then combined, the association order differs from a single-node pass. Measured: 11185 of 20000 random top-8 draws (56%) sum differently — by exactly 1 ULP — under the two orders. The M2 gate would fail on arithmetic alone, with every per-tensor check still green. | The reduction must be canonical by construction: nodes exchange their per-expert contributions and every node accumulates in fixed expert-id order, so a single-node pass and an N-node pass perform the same additions in the same sequence. The payload grows to top-k × 4 KB (≤ 40 KB) — affordable, since the transport is latency-bound (Testbed). This is a P0/P3 decision, not an implementation detail. |
| R11 | bf16 router logits flip the top-k set. I3 says discrete decisions must match exactly and keeps routers at "bf16 or higher". Measured over 20,000 random 256-way routers with round-to-nearest-even bf16: 922 (4.61%) changed the top-8 index set. A concrete pair: fp32 logits 5.363673687 and 5.385571480 both round to bf16 5.375, so the ordering that separates them is gone. One flipped index diverges the output completely while per-tensor MSE still looks healthy. (Correction, 2026-09-15: an earlier version of this row quoted a pair at 5.424064/5.413122 — those came from mantissa truncation, which is not what a bf16 cast does. The rate is unchanged in substance; the example was wrong and is replaced.) | Resolved as D5: router logits, normalization, comparison and top-k run in fp32, with an explicit ascending-expert-id tie-break applied identically on both sides; bf16 or quantized router weights are fine because the accumulation is fp32. The I3 assertion is a separate test from any numeric tolerance and needs a fixture whose logits sit within 1 ULP of each other, or it cannot fail — test_bf16_rounding_reproduces_the_router_hazard is that fixture. |
| R12 | M0's model is not the conventional one the milestone was designed around. D1 chose Qwen3.5-2B, whose first layer is a chunked Gated DeltaNet: the bit-exactness harness would be first exercised against a chunked delta rule with cumulative decay, a depthwise causal conv and a gated norm — the kernel family the brief assigns to M4/M5 — and, if 4-bit lands in the same step, against a new quantization path at the same time. Two independent error sources in the first proof of the harness is how a harness gets quietly distrusted. | The milestone is staged (M0a harness + conventional model, M0b bf16 on the real architecture, M0c 4-bit) so exactly one variable moves at a time; the work is not wasted, because M1's Qwen3.6-35B-A3B is 30-of-40 layers of the same Gated DeltaNet family, and the sister project already runs that family single-node as a comparison point. |
| R13 | PyTorch's matmul order is not reproducible by any other implementation. Measured on a real layer's shapes (8×1024 by 512×1024, fp32): an explicitly ordered accumulation differs from torch's in 3544 of 4096 outputs (mean 17 ULP, max 22587 ULP), while both sit exactly 3.148e-07 from an fp64 computation of the same product. The difference is summation order, not accuracy, and the order belongs to torch's BLAS kernels rather than to the model. A gate defined as "bit-match torch" would fail forever, for a reason that has nothing to do with the engine. | Resolved in D3 by splitting the references: tools/ordered_reference.py states the order of every sum and is the bit-exactness target, while the torch trace is the semantic oracle — exact discrete decisions (I3) and per-tensor closeness. The engine's kernels must accumulate in the contract's order (no split-K, no FMA, no reassociation), and I2 compares our implementation to itself. Confirmed on the real model: the ordered forward's logits argmax agrees with torch at every position while the last bits differ from the first matmul onward. |
Decisions that are not yet made. Each one is answered by an ADR or a task above, and the answer is recorded here even when it is "not yet decided".
| ID | Question | Answered by |
|---|---|---|
| Q1 | Which dense model is M0's target, and which implementation produces its golden traces? | DC-020, DC-021 |
| Q2 | Is the IR a shared component with the sister runtime, or a new one built beside it? | DC-010 |
| Q3 | How does the 1-node M1 baseline run the sharded model for a bit-identical comparison — one shard holding every expert, or a separate unsharded path? | DC-006, DC-035 |
| Q4 | What exactly stays resident on each 8 GB node, and what is streamed? | DC-011, DC-032 |
| Q5 | Which transport is mandatory for M2/M3, and does the Thunderbolt bridge need elevated privileges? (Measured: 1 Gbit Ethernet 0.49–0.64 ms vs mesh VPN 1.4–1.8 ms; no bridge is configured — see R4 and DC-017.) | DC-008 |
| Q6 | Does "unlimited nodes" mean N is bounded by experts-per-layer divided by k, or is expert duplication allowed? | DC-011, DC-050 |
| Q7 | Where do golden traces live — in the repository, or in external storage? Trace size decides it. | DC-007 |
| Q8 | Is Core ML / Neural Engine prefill part of a node's engine, and if so where is the sidecar produced and verified? (The conversion toolchain is installed — DC-016 — but installing it decides nothing.) | DC-033, DC-061 |
| Q9 | Why 2,218 expert requests for a five-token prefill? That is 1,109 fetches, or 221.8 per token, against 8 experts × 40 layers = 320 selections per token. De-duplicating selections within a layer would land below 320, but that is an inference: settle it by counting distinct experts per layer in the mixture and comparing | DC-034 |
Ideas that are not scheduled and not promised. They are written down so they are not
lost, and they are not tasks until they get a DC-nnn id.
- A
docs/tree in the repository for runbooks and audit registers, mirroring the sister project's split between wiki (user/engineering documentation) anddocs/(runbooks). - Package the cluster launcher as a single script that reports per-node readiness before a run, in the spirit of the sister project's launcher.
- A public benchmark page once there is a number worth publishing — with the prompt set and the hardware named, or not published at all.
- Report the sister project's
SECURITY.mddefect upstream: its private vulnerability reporting link points atgithub.com/drumih/turbo-fieldfare, not at TinyTitan. Found on 2026-09-15 while writing this repository's own policy; it belongs in the sibling tracker or an upstream issue (DC-085).
Source · Issues · Discussions · Sister project: TinyTitan
Project
Design
Development
Sister project: TinyTitan