Repository navigation
Target Models
The three models this project exists to run, verified against the vendors' own
configuration files on 2026-09-15 rather than taken from a brief. Every figure below
came out of the checkpoint's config.json at the revision named next to it; the
revision is the provenance anchor that I6 requires in the converted artifact.
Nothing here is implemented. The op-level semantics — norm placement, epsilon, RoPE application order, the n-gram hash — are not in a config file, and must be read out of the reference implementations before a single kernel is written.
| Qwen3.6-35B-A3B | DeepSeek-V4-Flash | Qwen3.8-Flash-Next | |
|---|---|---|---|
| Revision | 995ad96e… |
60d8d707… |
de4b8e4d… |
model_type |
qwen3_5_moe |
deepseek_v4 |
qwen4_exp |
| Layers | 40 | 43 | 48 |
| — attention mix | 30 linear + 10 full | hybrid CSA + HCA | 36 linear + 12 full |
| Hidden | 2048 | 4096 | 2560 |
| Routed experts | 256 | 256 | 512 |
| Top-k | 8 | 6 | 10 |
| Shared expert width | 512 | 1 expert | 640 |
| MoE inner width | 512 | 2048 | 640 |
| Context | 262,144 | 1,048,576 | 262,144 |
| Vocabulary | 248,320 | 129,280 | 248,320 |
| Expert dtype | bf16 | FP4 | bf16 |
| Dense dtype | bf16 | FP8 e4m3, 128×128 blocks | bf16 |
| Extra | output gate, MTP head | 3 hash-routed layers, indexer (64 heads, top-512), noaux_tc, 1 KV head, MTP |
hyper-connections (4 branches, rank 320), sparse indexer (top-2048), tri-gram table (20M entries × 8 heads), MTP |
Three differences matter, and each changes the plan rather than the wording:
- Qwen3.6-35B-A3B is not a conventional model. 30 of its 40 layers are linear attention (the Gated-DeltaNet line), with a full-attention layer every fourth one, plus an attention output gate and an MTP head. The brief treats M1 as the cheap validation step; as configured, M1 already needs the recurrence kernels the brief assigns to M5. Either M1 costs more than planned, or the M1 validation model should be a genuinely dense-attention MoE.
-
DeepSeek-V4-Flash has mechanisms the brief does not mention: 3 hash-routed layers,
a 64-head indexer with a top-512 budget,
noaux_tcrouting, a single KV head, vocab 129,280, and FP8 dense weights in 128×128 blocks — that block geometry is what any FP8 transcoder has to match. -
The Qwen3.8 n-gram table is a tri-gram table (
ngram_size: 3) of 20,000,000 entries with 8 heads per n-gram andsplit_ngram_parts: 128. It is not "bigram and trigram".
Everything else in the brief checked out: 43 layers, 256 experts, top-6 + 1 shared, FP4 experts, FP8 dense, 1M context, MIT licence for DeepSeek; 48 layers, 512 experts, top-10
- 1 shared, MTP and the 3-to-1 linear/full rhythm for Qwen3.8.
Expert size is 3 × hidden × moe_inner; at 4-bit that is 0.5 bytes per parameter, at FP4
the same. "All miss" means no expert is cached — the worst case the cache exists to
avoid.
| Expert size | Bytes read per token, all miss | At the measured 1161 MB/s sequential | At the measured 108 MB/s QD1 random | |
|---|---|---|---|---|
| Qwen3.6-35B-A3B | 3.1 M params ≈ 1.6 MB | 8 × 1.6 MB × 40 = 503 MB | 0.43 s | 4.7 s |
| DeepSeek-V4-Flash | 25.2 M params ≈ 12.6 MB | 6 × 12.6 MB × 43 = 3.25 GB | 2.8 s | 30 s |
| Qwen3.8-Flash-Next | 4.9 M params ≈ 2.5 MB | 10 × 2.5 MB × 48 = 1.18 GB | 1.0 s | 11 s |
The last two columns are the whole argument for expert caching, expert reordering and a deep read queue: the same bytes cost 9–11x more when they are random slabs at queue depth 1 than when they stream sequentially. A cache-hit-rate measurement is therefore not a nice-to-have for M1's gate; it is the number that decides whether the model runs at all.
One all-reduce per MoE layer, 4 KB payload, on the measured 1 Gbit segment (0.52 ms ICMP,
0.64 ms for 4 KB over TCP). A ring needs 2(N-1) sequential steps; recursive doubling
needs log2(N) parallel rounds. Per token, over the whole stack:
| Topology | N=2 | N=4 | N=8 |
|---|---|---|---|
| Ring, 40 MoE layers | 26 ms | 77 ms | 180 ms |
| Recursive doubling, 40 layers | 26 ms | 26 ms | 38 ms |
| Ring, 48 MoE layers | 31 ms | 91 ms | 215 ms |
| Recursive doubling, 48 layers | 31 ms | 31 ms | 46 ms |
At a decode budget of 150–250 ms per token a ring at four nodes would spend a third to a half of it on synchronisation alone. Determinism needs a fixed reduction order, not a ring — see R10 and R11 in the Project Tracker.
M0 needs a small dense model, and the brief's "~1B" has no exact member in the Qwen3
line: the candidates are Qwen/Qwen3-1.7B and Qwen/Qwen3-0.6B, both qwen3
(28 layers, RMSNorm + RoPE + GQA + SwiGLU, tied embeddings). Qwen3-1.7B is the working
recommendation — conventional, first-class reference, and its tied head avoids the untied
case the runtime already tripped over. It is pinned for the reference work at
revision 70d244cc86ccca08cf5af4e1e306ecf908b1ad5e; the choice itself is open as
DC-020/D1.
The op-level contract — every dtype boundary, the per-head QK-norm that is easy to miss,
the fp32 softmax island, the rotate-half RoPE with bf16 cos/sin, and the legacy
rope_theta key the v5 reference translates — is recorded in the repository at
docs/reference-qwen35-2b.md (for this model) and docs/reference-qwen3-dense.md (for
the conventional model that validates the harness), sourced from transformers v5.17.0
with each file's sha256. They exist so that no kernel has to guess — and, in the Qwen3.5
case, so that the Gated DeltaNet wiring is on the record before the milestone that has to
match it.
Source · Issues · Discussions · Sister project: TinyTitan
Project
Design
Development
Sister project: TinyTitan