Wiki review: remove the dependency framing; flag the plan pages as stale
Removed every non-historical 'sister project' reference (Home, Glossary, two Architecture rows, the
Roadmap objective, Testbed's dependency section) and the two tracker rows that presupposed adopting
another project's code - DC-085 and DC-123. ttd is built from scratch and is MIT, so there is no
upstream to track and no code to adopt.
Flagged rather than rewritten: Roadmap.md, the milestone pages and most tracker rows still describe
the retired DatacenterEngine plan, and Architecture.md's sharding section says expert parallelism
while what shipped divides layers. That is a planning decision, not a cleanup.
Tracker: D1-D7 recorded, M0 staged, the model verified before it is built on
DC-020 closes with D1 and DC-002 closes with D7. M0 becomes three stages with
one variable each, because D1's model is a chunked Gated DeltaNet rather than
the conventional transformer the milestone was designed around (R12) — and
because the 4-bit path should not land in the same step as the first
bit-exactness proof.
The Target models page and the Roadmap carry the same staging, and the
decisions themselves live in the repository at docs/m0-decisions.md so the
reasoning survives outside this wiki. What is still missing from the Gated
DeltaNet contract — the intra-chunk algorithm, the conv boundary handling, the
l2norm reduction, the full-attention module — is listed explicitly rather than
left to be guessed.
Tracker: the M0 reference contract is written (DC-027)
DC-027 closes with docs/m0-reference-contract.md — the official implementation
read at a pinned tag, every dtype boundary recorded with the file and line it
came from, and the handful of details that decide whether M0 matches the
reference or merely looks close to it: the per-head QK-norm before RoPE, the
norm casting to bf16 before the weight multiply, the fp32 softmax island, the
rotate-half RoPE with bf16 cos/sin, the tied head, and the legacy rope_theta
key the v5 reference translates but the checkpoint still ships.
The Target models page gains the M0 candidate: Qwen3-1.7B, pinned at revision
70d244cc for the reference work, with Qwen3-0.6B as the smaller option — and
the note that pinning a candidate is not choosing it. DC-020 and D1 stay open,
because "~1B" has no exact member in the Qwen3 line and that is the operator's
call.
Also recorded: torch publishes cp314 macOS arm64 wheels, so the reference
harness runs on the project's Python 3.14 standard rather than forcing a
downgrade — a question that was open when the Python standard was set.
Baselines measured, target models verified, and two arithmetic hazards
DC-019: the numbers the plan depends on, measured rather than assumed, plus
the three target models read out of their own checkpoints.
Measured (Testbed, "Measured baselines"):
- 4 KB round trip on the 1 Gbit segment: 645 us over TCP, 687 us over UDP,
against a 515 us empty ICMP round trip. The transport is base-RTT-bound, not
bandwidth-bound, so the 4 KB payload is nearly free and UDP buys nothing.
- Internal SSD: 1161 MB/s sequential read, but 108 MB/s for random 16 KB reads
at queue depth 1 — 9-11x slower. That gap is the entire case for a deep read
queue and expert caching, and it is the number the bytes-per-token maths has
to use.
- ~2.5 GB reclaimable memory on an idle 8 GB node; no external NVMe attached on
any node, so the assumed 3 GB/s has no hardware behind it yet.
Two probes were thrown away before this: the first reported 13 us round trips
(a dead echo server returning EOF instantly) and the second reported 13 GB/s
reads (a 2 GB file that fit in the 8 GB machine's cache). Both are recorded as
discarded, not published — a plausible wrong number is worse than none.
Target models (new page): Qwen3.6-35B-A3B, DeepSeek-V4-Flash and
Qwen3.8-Flash-Next read from their own config.json at a named revision, with
what the brief got right and the three things it does not say — Qwen3.6 is 30
linear + 10 full attention layers, not conventional; DeepSeek has 3 hash-routed
layers, a 64-head indexer and 128x128 FP8 blocks; the Qwen n-gram table is a
tri-gram table of 20M entries.
Two hazards, both demonstrated with numbers rather than argued:
- R10: bit-identity across shard counts does not follow from "accumulate in
fp32". Pre-summing each node's own experts and combining the partials is a
different association order: 11185 of 20000 random top-8 draws (56%) sum
differently, by exactly 1 ULP. The M2 gate would fail on arithmetic alone.
Nodes must exchange per-expert contributions and accumulate in fixed
expert-id order — determinism needs a fixed order, not a ring.
- R11: bf16 router logits flip the top-k set. I3 keeps routers at "bf16 or
higher"; over random 256-way routers the top-8 index set changed within 22
draws (expert 102 at 5.424064 rounded to expert 85's 5.413122). Routers go
through fp32 logits and fp32 top-k, with a documented tie-break.
The sync-cost table on the Target models page turns R4's latency into a per
token budget: on the measured link a ring is 77 ms/token at four nodes over 40
MoE layers, against 26 ms for recursive doubling.