Skip to content

History / Target Models

Revisions

  • Wiki review: remove the dependency framing; flag the plan pages as stale Removed every non-historical 'sister project' reference (Home, Glossary, two Architecture rows, the Roadmap objective, Testbed's dependency section) and the two tracker rows that presupposed adopting another project's code - DC-085 and DC-123. ttd is built from scratch and is MIT, so there is no upstream to track and no code to adopt. Flagged rather than rewritten: Roadmap.md, the milestone pages and most tracker rows still describe the retired DatacenterEngine plan, and Architecture.md's sharding section says expert parallelism while what shipped divides layers. That is a planning decision, not a cleanup.

    @Pummelchen Pummelchen committed Sep 19, 2026
  • Tracker: D1-D7 recorded, M0 staged, the model verified before it is built on DC-020 closes with D1 and DC-002 closes with D7. M0 becomes three stages with one variable each, because D1's model is a chunked Gated DeltaNet rather than the conventional transformer the milestone was designed around (R12) — and because the 4-bit path should not land in the same step as the first bit-exactness proof. The Target models page and the Roadmap carry the same staging, and the decisions themselves live in the repository at docs/m0-decisions.md so the reasoning survives outside this wiki. What is still missing from the Gated DeltaNet contract — the intra-chunk algorithm, the conv boundary handling, the l2norm reduction, the full-attention module — is listed explicitly rather than left to be guessed.

    @Pummelchen Pummelchen committed Sep 15, 2026
  • Tracker: the M0 reference contract is written (DC-027) DC-027 closes with docs/m0-reference-contract.md — the official implementation read at a pinned tag, every dtype boundary recorded with the file and line it came from, and the handful of details that decide whether M0 matches the reference or merely looks close to it: the per-head QK-norm before RoPE, the norm casting to bf16 before the weight multiply, the fp32 softmax island, the rotate-half RoPE with bf16 cos/sin, the tied head, and the legacy rope_theta key the v5 reference translates but the checkpoint still ships. The Target models page gains the M0 candidate: Qwen3-1.7B, pinned at revision 70d244cc for the reference work, with Qwen3-0.6B as the smaller option — and the note that pinning a candidate is not choosing it. DC-020 and D1 stay open, because "~1B" has no exact member in the Qwen3 line and that is the operator's call. Also recorded: torch publishes cp314 macOS arm64 wheels, so the reference harness runs on the project's Python 3.14 standard rather than forcing a downgrade — a question that was open when the Python standard was set.

    @Pummelchen Pummelchen committed Sep 15, 2026
  • Baselines measured, target models verified, and two arithmetic hazards DC-019: the numbers the plan depends on, measured rather than assumed, plus the three target models read out of their own checkpoints. Measured (Testbed, "Measured baselines"): - 4 KB round trip on the 1 Gbit segment: 645 us over TCP, 687 us over UDP, against a 515 us empty ICMP round trip. The transport is base-RTT-bound, not bandwidth-bound, so the 4 KB payload is nearly free and UDP buys nothing. - Internal SSD: 1161 MB/s sequential read, but 108 MB/s for random 16 KB reads at queue depth 1 — 9-11x slower. That gap is the entire case for a deep read queue and expert caching, and it is the number the bytes-per-token maths has to use. - ~2.5 GB reclaimable memory on an idle 8 GB node; no external NVMe attached on any node, so the assumed 3 GB/s has no hardware behind it yet. Two probes were thrown away before this: the first reported 13 us round trips (a dead echo server returning EOF instantly) and the second reported 13 GB/s reads (a 2 GB file that fit in the 8 GB machine's cache). Both are recorded as discarded, not published — a plausible wrong number is worse than none. Target models (new page): Qwen3.6-35B-A3B, DeepSeek-V4-Flash and Qwen3.8-Flash-Next read from their own config.json at a named revision, with what the brief got right and the three things it does not say — Qwen3.6 is 30 linear + 10 full attention layers, not conventional; DeepSeek has 3 hash-routed layers, a 64-head indexer and 128x128 FP8 blocks; the Qwen n-gram table is a tri-gram table of 20M entries. Two hazards, both demonstrated with numbers rather than argued: - R10: bit-identity across shard counts does not follow from "accumulate in fp32". Pre-summing each node's own experts and combining the partials is a different association order: 11185 of 20000 random top-8 draws (56%) sum differently, by exactly 1 ULP. The M2 gate would fail on arithmetic alone. Nodes must exchange per-expert contributions and accumulate in fixed expert-id order — determinism needs a fixed order, not a ring. - R11: bf16 router logits flip the top-k set. I3 keeps routers at "bf16 or higher"; over random 256-way routers the top-8 index set changed within 22 draws (expert 102 at 5.424064 rounded to expert 85's 5.413122). Routers go through fp32 logits and fp32 top-k, with a documented tie-break. The sync-cost table on the Target models page turns R4's latency into a per token budget: on the measured link a ring is 77 ms/token at four nodes over 40 MoE layers, against 26 ms for recursive doubling.

    @Pummelchen Pummelchen committed Sep 15, 2026