Skip to content

Target Models

Pummelchen edited this page Sep 19, 2026 · 4 revisions
TinyTitan Datacenter

Target models

The three models this project exists to run, verified against the vendors' own configuration files on 2026-09-15 rather than taken from a brief. Every figure below came out of the checkpoint's config.json at the revision named next to it; the revision is the provenance anchor that I6 requires in the converted artifact.

Nothing here is implemented. The op-level semantics — norm placement, epsilon, RoPE application order, the n-gram hash — are not in a config file, and must be read out of the reference implementations before a single kernel is written.

Verified configuration

Qwen3.6-35B-A3B DeepSeek-V4-Flash Qwen3.8-Flash-Next
Revision 995ad96e… 60d8d707… de4b8e4d…
model_type qwen3_5_moe deepseek_v4 qwen4_exp
Layers 40 43 48
— attention mix 30 linear + 10 full hybrid CSA + HCA 36 linear + 12 full
Hidden 2048 4096 2560
Routed experts 256 256 512
Top-k 8 6 10
Shared expert width 512 1 expert 640
MoE inner width 512 2048 640
Context 262,144 1,048,576 262,144
Vocabulary 248,320 129,280 248,320
Expert dtype bf16 FP4 bf16
Dense dtype bf16 FP8 e4m3, 128×128 blocks bf16
Extra output gate, MTP head 3 hash-routed layers, indexer (64 heads, top-512), noaux_tc, 1 KV head, MTP hyper-connections (4 branches, rank 320), sparse indexer (top-2048), tri-gram table (20M entries × 8 heads), MTP

What the brief said, and what is actually there

Three differences matter, and each changes the plan rather than the wording:

  1. Qwen3.6-35B-A3B is not a conventional model. 30 of its 40 layers are linear attention (the Gated-DeltaNet line), with a full-attention layer every fourth one, plus an attention output gate and an MTP head. The brief treats M1 as the cheap validation step; as configured, M1 already needs the recurrence kernels the brief assigns to M5. Either M1 costs more than planned, or the M1 validation model should be a genuinely dense-attention MoE.
  2. DeepSeek-V4-Flash has mechanisms the brief does not mention: 3 hash-routed layers, a 64-head indexer with a top-512 budget, noaux_tc routing, a single KV head, vocab 129,280, and FP8 dense weights in 128×128 blocks — that block geometry is what any FP8 transcoder has to match.
  3. The Qwen3.8 n-gram table is a tri-gram table (ngram_size: 3) of 20,000,000 entries with 8 heads per n-gram and split_ngram_parts: 128. It is not "bigram and trigram".

Everything else in the brief checked out: 43 layers, 256 experts, top-6 + 1 shared, FP4 experts, FP8 dense, 1M context, MIT licence for DeepSeek; 48 layers, 512 experts, top-10

  • 1 shared, MTP and the 3-to-1 linear/full rhythm for Qwen3.8.

Derived from the verified configs

Expert size is 3 × hidden × moe_inner; at 4-bit that is 0.5 bytes per parameter, at FP4 the same. "All miss" means no expert is cached — the worst case the cache exists to avoid.

Expert size Bytes read per token, all miss At the measured 1161 MB/s sequential At the measured 108 MB/s QD1 random
Qwen3.6-35B-A3B 3.1 M params ≈ 1.6 MB 8 × 1.6 MB × 40 = 503 MB 0.43 s 4.7 s
DeepSeek-V4-Flash 25.2 M params ≈ 12.6 MB 6 × 12.6 MB × 43 = 3.25 GB 2.8 s 30 s
Qwen3.8-Flash-Next 4.9 M params ≈ 2.5 MB 10 × 2.5 MB × 48 = 1.18 GB 1.0 s 11 s

The last two columns are the whole argument for expert caching, expert reordering and a deep read queue: the same bytes cost 9–11x more when they are random slabs at queue depth 1 than when they stream sequentially. A cache-hit-rate measurement is therefore not a nice-to-have for M1's gate; it is the number that decides whether the model runs at all.

Cross-node synchronisation, at the measured latency

One all-reduce per MoE layer, 4 KB payload, on the measured 1 Gbit segment (0.52 ms ICMP, 0.64 ms for 4 KB over TCP). A ring needs 2(N-1) sequential steps; recursive doubling needs log2(N) parallel rounds. Per token, over the whole stack:

Topology N=2 N=4 N=8
Ring, 40 MoE layers 26 ms 77 ms 180 ms
Recursive doubling, 40 layers 26 ms 26 ms 38 ms
Ring, 48 MoE layers 31 ms 91 ms 215 ms
Recursive doubling, 48 layers 31 ms 31 ms 46 ms

At a decode budget of 150–250 ms per token a ring at four nodes would spend a third to a half of it on synchronisation alone. Determinism needs a fixed reduction order, not a ring — see R10 and R11 in the Project Tracker.

The M0 dense model (candidate, not yet chosen)

M0 needs a small dense model, and the brief's "~1B" has no exact member in the Qwen3 line: the candidates are Qwen/Qwen3-1.7B and Qwen/Qwen3-0.6B, both qwen3 (28 layers, RMSNorm + RoPE + GQA + SwiGLU, tied embeddings). Qwen3-1.7B is the working recommendation — conventional, first-class reference, and its tied head avoids the untied case the runtime already tripped over. It is pinned for the reference work at revision 70d244cc86ccca08cf5af4e1e306ecf908b1ad5e; the choice itself is open as DC-020/D1.

The op-level contract — every dtype boundary, the per-head QK-norm that is easy to miss, the fp32 softmax island, the rotate-half RoPE with bf16 cos/sin, and the legacy rope_theta key the v5 reference translates — is recorded in the repository at docs/reference-qwen35-2b.md (for this model) and docs/reference-qwen3-dense.md (for the conventional model that validates the harness), sourced from transformers v5.17.0 with each file's sha256. They exist so that no kernel has to guess — and, in the Qwen3.5 case, so that the Gated DeltaNet wiring is on the record before the milestone that has to match it.

TinyTitan Datacenter

Project

Design

Development

Sister project: TinyTitan

Clone this wiki locally