Repository navigation
Roadmap
This page is the plan: what each milestone is, what it delivers, and the gate it has to pass. Task-level status lives in the Project Tracker, while the milestone state lives here. Terms are defined in the Glossary.
Important
P1 (M0), P2 (M1) and P3 (M2) are complete and gated; P4 (M3) is under way — its functional half is
demonstrated (four machines, one bit-identical trace) and its throughput gate needs a quiet farm. M0's gate
passed on Qwen/Qwen3.5-2B — three frozen prompts, 40,683,520 bytes identical to the
contract, every discrete decision matching — and M0c measured what 4-bit costs. M1's gate
passes on the real 35 B model: the trace is byte-identical to the contract
(digest b8c976c5e7ba8816…), generation is 0.108 tok/s cached, peak memory is
348.6 MB. M2's gate passes across two machines: the real 35 B model over a
256-expert plan produced the same digest as the single-node baseline
(b0d382dbabf36df0…), 83 tensors and 40 discrete decisions checked on both nodes. What
these leave open is the tok/s measurements themselves, which need a quiet farm. D12, the design
question about the expert banks, was decided by measurement (D31: sized from a budget, one slot,
because the hit rate is 0 at every size). DC-108 (a Python contract that reads an install) and DC-109
(a sharded generation CLI) are both closed: the install now verifies from Python, structure, tiling,
policy and every payload digest, and a two-machine cached generation produced the reference's tokens
and digest. The gates below are the definition of done this project holds
itself to, not a schedule, and a phase ends when its gate passes with the evidence
recorded, not when the code exists. Live state:
Project Tracker.
| Phase | P4 — M3 (4 nodes). P0 in progress; M0, M1, M2 gates passed |
| This node | node4 — four identical Mac mini M2, 8 GB, macOS 27.0, Xcode 27, Python 3.14.7 |
| Farm link | wired, on a switch. 1 Gbit Ethernet measured 0.49–0.64 ms round trip; the mesh VPN path measures 1.4–1.8 ms and is what node names resolve to |
| Tests |
266 Swift, 428 Python, nothing skipped — the Metal kernel tests run on the node's GPU since D34
|
| Toolchain |
Swift 6.4 on Xcode 27, swift-tools-version:6.4, enforced by tools/check_toolchain.py
|
| Open tasks | 27 — 17 open, 1 blocked, 9 parked, in the Project Tracker; everything closed, and the dated record, are on News |
Each milestone is one phase of the project. A phase ends when its gate passes on real hardware, with the evidence recorded — not when the code exists.
| Phase | Milestone | Objective | Gate (definition of done) |
|---|---|---|---|
| P0 | Foundation | Decisions, harnesses and conventions that every later phase depends on | The ADRs, the trace/diff harness spec and the repository conventions are written and reviewed |
| P1 | M0 | Single node, small dense model, bf16, through the IR | Bit-matches the reference golden traces |
| P2 | M1 | Qwen3.6-35B-A3B, single node, 4-bit, SSD-streamed | Correct output and a recorded tok/s baseline on a frozen prompt set |
| P3 | M2 | 2 nodes, expert-parallel, dense backbone replicated | Bit-identical to the M1 baseline — router top-k index sets included |
| P4 | M3 | 4 nodes | ≥3x the M1 tok/s on the same frozen prompt set |
| P5 | M4 | DeepSeek-V4.1-Flash | Matches the reference at 128K context |
| P6 | M5 | Qwen3.8-Flash-Next | Same gate protocol as M4 |
M2 is the real gate for the project: it is the point at which sharding is proven either to be transparent (bit-identical) or to be a source of silent numeric drift.
Objective. Fix the things that are expensive to change later: the reduction contract, the trace format, the transport, the IR, and the shard policy format.
Deliverables
- Architecture decision records: deterministic reduction contract, transport, all-reduce wire protocol, model IR schema, shard policy data format.
- The golden-trace and diff harness specification — what is captured, where, in what format, and what tolerance (zero) is applied.
- The bit-reproducibility test protocol: how an N-node result is compared with the 1-node result, including the comparison of discrete index sets.
- Repository conventions, licence/provenance review, and CI that can enforce the cheap gates (Markdown links) before any code exists.
Gate. Every decision above is written down, with the alternatives and the reason for rejecting them, and the harness specification is concrete enough to implement against.
Objective. Prove the harness before the model, and prove the model one variable at a time. D1 selected the dense Qwen3.5-2B — which is dense in the mixture sense (no routed experts) but not conventional: 24 layers, 18 of them Gated DeltaNet. Landing that, the 4-bit path and a brand-new harness in one step would confound three error sources, so M0 is staged:
| Stage | Model | What it proves |
|---|---|---|
| M0a | a synthetic fixture set plus Qwen3-1.7B (conventional) | the capture and diff harness works and can fail |
| M0b | Qwen3.5-2B, bf16 | the real architecture bit-exact, conventional path first, the Gated DeltaNet after |
| M0c | Qwen3.5-2B, 4-bit | the quantization/transcode path as a separate variable |
Deliverables
- A pinned reference runner that emits golden traces, computing in fp32 from bf16-stored weights, one layer resident at a time (D3, D6 — an fp32 2 B model does not fit in 8 GB).
- The IR importer for that model — a pure name-to-role mapping, text tower only.
- A minimal single-node forward pass on Metal that runs the model through the IR.
- The trace-capture harness and the diff harness, with first-divergence localisation and a separate assertion for discrete decisions (I3).
Gate. The runtime's traces are bit-identical to the reference traces for the pinned model and prompt set, at each stage, with the diff harness reporting the exact first divergent tensor when they are not — and the harness's own seeded-failure tests passing first, because a harness that cannot fail proves nothing.
What it does not prove. Expert routing, sharding, streaming, or cluster numerics.
Objective. Reach the single-node streaming capability this project targets inside this codebase, on the model family that M2 will shard.
Deliverables
- Importer for the Qwen3.6-35B-A3B family.
- The quantization/transcode path (native FP8/FP4 checkpoints are transcoded, not requantized).
- SSD expert streaming with a bounded per-node expert cache.
- Metal kernels for attention, router, MoE and head.
- A frozen prompt set and a recorded throughput baseline.
Gate. Correct output on the frozen prompt set, plus a recorded tok/s baseline. The baseline and the prompt set are then frozen — they are the comparison target for M2 and M3, so changing either afterwards invalidates the later gates and must be recorded in the tracker.
Objective. Prove expert parallelism is numerically transparent.
Deliverables
- The shard planner: disjoint 1/N expert assignment, dense backbone replicated.
- The all-reduce implementation over the chosen transport: fp32 accumulation, fixed reduction order, one reduction per MoE layer.
- Cluster bring-up: discovery, handshake, config/version hash agreement, fail fast.
- Node-failure and timeout semantics for a single-user run.
- The bit-identity gate harness.
Gate. Bit-identical to the M1 baseline — the generated text and the per-layer tensors match byte for byte, and the discrete decisions match exactly:
- router top-k index sets per layer and per token,
- sparse-attention block selections, once a model with sparse attention is in scope.
A green per-tensor MSE check with a differing top-k index set is a failure, not a pass: after the first flipped index the output diverges completely while every smoothness metric still looks healthy.
Objective. Turn a proof into a speedup.
Deliverables
- The shard planner generalised to N nodes of the same model type.
- A measured per-token synchronisation budget: all-reduce latency per MoE layer against the expert-read time it overlaps.
- Per-node cache and prefetch tuning at four nodes.
- Per-node observability: tokens/s, bytes read per token, all-reduce time, cache hit rate.
Gate. ≥3x the frozen M1 tok/s on the same prompt set, with the bit-identity gate from M2 still passing at four nodes.
Objective. Absorb a frontier family whose experts are natively FP4 and whose dense weights are FP8.
Deliverables
- The DeepSeek importer, with transcoding rather than requantization.
- Compressed / sparse hybrid attention and constant-size recurrent state.
- MTP speculative decoding parity.
- Sparse-attention block-selection exactness checks.
Gate. Matches the reference at 128K context, on the same exactness protocol as M2 (tensors bit-exact where the reference is deterministic; discrete selections exact).
Objective. The second frontier family, and the largest shape in scope.
Deliverables
- Importer, reference match, and the sparse n-gram lookup table path.
- Expert-parallel scaling on the 125B-A6B shape.
Gate. Same protocol as M4. An IR built around this shape should absorb the next generation with importer changes only — if it does not, that is a design finding, not a model problem.
- Gates are not renegotiable in private. If a gate turns out to be unattainable, the finding and the redefinition are recorded in the Project Tracker with the measurement that forced it. Silently weakening a gate is the one failure this project cannot afford, because every later claim is built on it.
- Cheap before expensive. M0 exists so the harness is trusted before the model is. The harness is what detects a conversion bug versus a scheduling artifact — on these models you will need that distinction.
- Frozen baselines. M1's prompt set and throughput baseline are frozen artifacts. Re-recording them is a documented event with a reason, never a convenience.
- One sequence, one user. Throughput work that only pays off under batching is out of scope; if a change does not help a single stream, it is not this project.
Source · Issues · Discussions · Sister project: TinyTitan
Project
Design
Development
Sister project: TinyTitan