Repository navigation
Glossary
Pummelchen edited this page Sep 19, 2026
·
2 revisions
Terms as they are used on these pages. Where a term means something project-specific, that is stated; where the project has not fixed a definition yet, it says so.
| Term | Meaning |
|---|---|
| MoE (mixture of experts) | A model whose feed-forward layers are split into many experts, of which a few are selected per token. Only the selected ones are computed — and, in this project, only those are read from disk. |
| Routed expert | One of the many feed-forward experts a router can choose. Streamed from SSD rather than kept resident. |
| Shared expert | An expert every token uses. It is always needed, so it stays resident with the dense backbone. |
| Top-k routing | The router picks the k highest-scoring experts per token per layer. The set of indices it picks is a discrete decision, not a numeric value — see invariant 2 on the Architecture page. |
| Dense backbone | Everything that is not a routed expert: attention, the router, the shared expert and the head. Replicated on every node under expert parallelism. |
| Dense model | A model with no routed experts at all. M0 uses one deliberately, so the harness can be proven without routing in the way. |
| Sparse / compressed attention | Attention variants that attend to a selected subset of positions or blocks instead of the whole context. The selection is discrete and must match the reference exactly. |
| Constant-size recurrent state | A state that does not grow with context length, used by newer hybrid attention designs. |
| MTP (multi-token prediction) | A speculative decoding scheme that drafts several tokens ahead for the main model to verify. |
| N-gram lookup table | A large precomputed table used by some models to shortcut predictable spans. It is big, and how it is shared across nodes is an open question. |
| FP4 / FP8 / bf16 | Weight and activation number formats. Frontier MoE models ship experts in FP4 or FP8 natively; bf16 is the format M0 works in. |
| Quantization | Packing weights into fewer bits. This project prefers the vendor's own quantized weights over re-quantizing. |
| Transcoding | Converting a quantized representation into the runtime's layout without dequantizing and re-quantizing. Invariant 4. |
| KV cache | The stored keys and values of previous tokens. It grows with context and competes for the same memory as the expert cache. |
| Term | Meaning |
|---|---|
| Expert parallelism | Sharding by expert: every node holds a disjoint 1/N slice of the routed experts and all nodes work on the same token. This project's core choice. |
| Pipeline parallelism | Sharding by layer: each node owns a contiguous block of layers. Rejected here, because with one sequence in flight only one node is busy per token. |
| All-reduce | Combining every node's partial MoE output into the result all nodes agree on. One per MoE layer, roughly 4 KB of payload, so latency dominates bandwidth. |
| Reduction order | The order in which partial sums are combined. Floating-point addition is not associative, so the order must be fixed for results to be reproducible at all. |
| Bit-identical | Byte-for-byte equal output. The M2 gate: two nodes must produce exactly what one node produced. |
| Discrete decision | A choice from a set — a top-k index set, a sparse-attention block selection — as opposed to a numeric tensor. Checked for exact equality, separately from numeric tolerances. |
| Golden trace | A recorded reference run: per-layer tensors, logits and discrete index sets, captured so a runtime's output can be diffed against it byte for byte. |
| Cold start | First token of a run, before any expert is cached. |
| Term | Meaning |
|---|---|
| IR (intermediate representation) | The declarative model description the runtime runs. It dispatches on tensor role, not tensor name, so architecture-specific knowledge lives only in importers. |
| Importer | The per-family module that maps a checkpoint's tensor names onto IR roles. Intended to be a pure mapping, roughly 500–800 lines. |
| Transform pass | An architecture-agnostic rewrite over the IR: fusing, repacking, expert reordering, quantization, sharding. |
| Shard policy | The data that says which node owns which experts, how the dense backbone is replicated, and which policy version is in use. |
| Bounded expert cache | The fixed memory budget an expert must fit in. Cache misses are SSD reads; the budget, not the page cache, is the working set. |
| SSD streaming | Reading expert weights from local storage at the moment they are selected, instead of holding the model in memory. |
| Term | Meaning |
|---|---|
| ttd | TinyTitan Datacenter: this project. Short name ttd. A dedicated repository, built from scratch and licensed under MIT. It is a the cluster. |
| TinyTitan Datacenter | This project — the distributed, expert-parallel engine. |
| Milestone (M0–M5) | A phase of the plan with a single gate. See the Roadmap. |
| Gate | The measurement that closes a milestone. A phase is not done because the code exists; it is done when its gate passes on real hardware. |
DC-nnn |
A stable task identifier in the Project Tracker. Never reused, never renumbered. |
| ADR | Architecture decision record: the choice, the alternatives and why they were rejected. |
Source · Issues · Discussions · Sister project: TinyTitan
Project
Design
Development
Sister project: TinyTitan