Skip to content

Roadmap

Pummelchen edited this page Sep 19, 2026 · 12 revisions
TinyTitan Datacenter

Roadmap

This page is the plan: what each milestone is, what it delivers, and the gate it has to pass. Task-level status lives in the Project Tracker, while the milestone state lives here. Terms are defined in the Glossary.

Important

P1 (M0), P2 (M1) and P3 (M2) are complete and gated; P4 (M3) is under way — its functional half is demonstrated (four machines, one bit-identical trace) and its throughput gate needs a quiet farm. M0's gate passed on Qwen/Qwen3.5-2B — three frozen prompts, 40,683,520 bytes identical to the contract, every discrete decision matching — and M0c measured what 4-bit costs. M1's gate passes on the real 35 B model: the trace is byte-identical to the contract (digest b8c976c5e7ba8816…), generation is 0.108 tok/s cached, peak memory is 348.6 MB. M2's gate passes across two machines: the real 35 B model over a 256-expert plan produced the same digest as the single-node baseline (b0d382dbabf36df0…), 83 tensors and 40 discrete decisions checked on both nodes. What these leave open is the tok/s measurements themselves, which need a quiet farm. D12, the design question about the expert banks, was decided by measurement (D31: sized from a budget, one slot, because the hit rate is 0 at every size). DC-108 (a Python contract that reads an install) and DC-109 (a sharded generation CLI) are both closed: the install now verifies from Python, structure, tiling, policy and every payload digest, and a two-machine cached generation produced the reference's tokens and digest. The gates below are the definition of done this project holds itself to, not a schedule, and a phase ends when its gate passes with the evidence recorded, not when the code exists. Live state: Project Tracker.

Current state

Phase P4 — M3 (4 nodes). P0 in progress; M0, M1, M2 gates passed
This node node4 — four identical Mac mini M2, 8 GB, macOS 27.0, Xcode 27, Python 3.14.7
Farm link wired, on a switch. 1 Gbit Ethernet measured 0.49–0.64 ms round trip; the mesh VPN path measures 1.4–1.8 ms and is what node names resolve to
Tests 266 Swift, 428 Python, nothing skipped — the Metal kernel tests run on the node's GPU since D34
Toolchain Swift 6.4 on Xcode 27, swift-tools-version:6.4, enforced by tools/check_toolchain.py
Open tasks 27 — 17 open, 1 blocked, 9 parked, in the Project Tracker; everything closed, and the dated record, are on News

Milestones, phases and gates

Each milestone is one phase of the project. A phase ends when its gate passes on real hardware, with the evidence recorded — not when the code exists.

Phase Milestone Objective Gate (definition of done)
P0 Foundation Decisions, harnesses and conventions that every later phase depends on The ADRs, the trace/diff harness spec and the repository conventions are written and reviewed
P1 M0 Single node, small dense model, bf16, through the IR Bit-matches the reference golden traces
P2 M1 Qwen3.6-35B-A3B, single node, 4-bit, SSD-streamed Correct output and a recorded tok/s baseline on a frozen prompt set
P3 M2 2 nodes, expert-parallel, dense backbone replicated Bit-identical to the M1 baseline — router top-k index sets included
P4 M3 4 nodes ≥3x the M1 tok/s on the same frozen prompt set
P5 M4 DeepSeek-V4.1-Flash Matches the reference at 128K context
P6 M5 Qwen3.8-Flash-Next Same gate protocol as M4

M2 is the real gate for the project: it is the point at which sharding is proven either to be transparent (bit-identical) or to be a source of silent numeric drift.

P0 — Foundation

Objective. Fix the things that are expensive to change later: the reduction contract, the trace format, the transport, the IR, and the shard policy format.

Deliverables

  • Architecture decision records: deterministic reduction contract, transport, all-reduce wire protocol, model IR schema, shard policy data format.
  • The golden-trace and diff harness specification — what is captured, where, in what format, and what tolerance (zero) is applied.
  • The bit-reproducibility test protocol: how an N-node result is compared with the 1-node result, including the comparison of discrete index sets.
  • Repository conventions, licence/provenance review, and CI that can enforce the cheap gates (Markdown links) before any code exists.

Gate. Every decision above is written down, with the alternatives and the reason for rejecting them, and the harness specification is concrete enough to implement against.

P1 — M0: staged — harness, then Qwen3.5-2B bf16, then 4-bit

Objective. Prove the harness before the model, and prove the model one variable at a time. D1 selected the dense Qwen3.5-2B — which is dense in the mixture sense (no routed experts) but not conventional: 24 layers, 18 of them Gated DeltaNet. Landing that, the 4-bit path and a brand-new harness in one step would confound three error sources, so M0 is staged:

Stage Model What it proves
M0a a synthetic fixture set plus Qwen3-1.7B (conventional) the capture and diff harness works and can fail
M0b Qwen3.5-2B, bf16 the real architecture bit-exact, conventional path first, the Gated DeltaNet after
M0c Qwen3.5-2B, 4-bit the quantization/transcode path as a separate variable

Deliverables

  • A pinned reference runner that emits golden traces, computing in fp32 from bf16-stored weights, one layer resident at a time (D3, D6 — an fp32 2 B model does not fit in 8 GB).
  • The IR importer for that model — a pure name-to-role mapping, text tower only.
  • A minimal single-node forward pass on Metal that runs the model through the IR.
  • The trace-capture harness and the diff harness, with first-divergence localisation and a separate assertion for discrete decisions (I3).

Gate. The runtime's traces are bit-identical to the reference traces for the pinned model and prompt set, at each stage, with the diff harness reporting the exact first divergent tensor when they are not — and the harness's own seeded-failure tests passing first, because a harness that cannot fail proves nothing.

What it does not prove. Expert routing, sharding, streaming, or cluster numerics.

P2 — M1: Qwen3.6-35B-A3B, single node, 4-bit, SSD-streamed

Objective. Reach the single-node streaming capability this project targets inside this codebase, on the model family that M2 will shard.

Deliverables

  • Importer for the Qwen3.6-35B-A3B family.
  • The quantization/transcode path (native FP8/FP4 checkpoints are transcoded, not requantized).
  • SSD expert streaming with a bounded per-node expert cache.
  • Metal kernels for attention, router, MoE and head.
  • A frozen prompt set and a recorded throughput baseline.

Gate. Correct output on the frozen prompt set, plus a recorded tok/s baseline. The baseline and the prompt set are then frozen — they are the comparison target for M2 and M3, so changing either afterwards invalidates the later gates and must be recorded in the tracker.

P3 — M2: 2 nodes, expert-parallel

Objective. Prove expert parallelism is numerically transparent.

Deliverables

  • The shard planner: disjoint 1/N expert assignment, dense backbone replicated.
  • The all-reduce implementation over the chosen transport: fp32 accumulation, fixed reduction order, one reduction per MoE layer.
  • Cluster bring-up: discovery, handshake, config/version hash agreement, fail fast.
  • Node-failure and timeout semantics for a single-user run.
  • The bit-identity gate harness.

Gate. Bit-identical to the M1 baseline — the generated text and the per-layer tensors match byte for byte, and the discrete decisions match exactly:

  • router top-k index sets per layer and per token,
  • sparse-attention block selections, once a model with sparse attention is in scope.

A green per-tensor MSE check with a differing top-k index set is a failure, not a pass: after the first flipped index the output diverges completely while every smoothness metric still looks healthy.

P4 — M3: 4 nodes

Objective. Turn a proof into a speedup.

Deliverables

  • The shard planner generalised to N nodes of the same model type.
  • A measured per-token synchronisation budget: all-reduce latency per MoE layer against the expert-read time it overlaps.
  • Per-node cache and prefetch tuning at four nodes.
  • Per-node observability: tokens/s, bytes read per token, all-reduce time, cache hit rate.

Gate. ≥3x the frozen M1 tok/s on the same prompt set, with the bit-identity gate from M2 still passing at four nodes.

P5 — M4: DeepSeek-V4.1-Flash

Objective. Absorb a frontier family whose experts are natively FP4 and whose dense weights are FP8.

Deliverables

  • The DeepSeek importer, with transcoding rather than requantization.
  • Compressed / sparse hybrid attention and constant-size recurrent state.
  • MTP speculative decoding parity.
  • Sparse-attention block-selection exactness checks.

Gate. Matches the reference at 128K context, on the same exactness protocol as M2 (tensors bit-exact where the reference is deterministic; discrete selections exact).

P6 — M5: Qwen3.8-Flash-Next

Objective. The second frontier family, and the largest shape in scope.

Deliverables

  • Importer, reference match, and the sparse n-gram lookup table path.
  • Expert-parallel scaling on the 125B-A6B shape.

Gate. Same protocol as M4. An IR built around this shape should absorb the next generation with importer changes only — if it does not, that is a design finding, not a model problem.

On ordering

  • Gates are not renegotiable in private. If a gate turns out to be unattainable, the finding and the redefinition are recorded in the Project Tracker with the measurement that forced it. Silently weakening a gate is the one failure this project cannot afford, because every later claim is built on it.
  • Cheap before expensive. M0 exists so the harness is trusted before the model is. The harness is what detects a conversion bug versus a scheduling artifact — on these models you will need that distinction.
  • Frozen baselines. M1's prompt set and throughput baseline are frozen artifacts. Re-recording them is a documented event with a reason, never a convenience.
  • One sequence, one user. Throughput work that only pays off under batching is out of scope; if a change does not help a single stream, it is not this project.
TinyTitan Datacenter

Project

Design

Development

Sister project: TinyTitan

Clone this wiki locally