Skip to content

Glossary

Pummelchen edited this page Sep 19, 2026 · 2 revisions
TinyTitan Datacenter

Glossary

Terms as they are used on these pages. Where a term means something project-specific, that is stated; where the project has not fixed a definition yet, it says so.

Model terms

Term Meaning
MoE (mixture of experts) A model whose feed-forward layers are split into many experts, of which a few are selected per token. Only the selected ones are computed — and, in this project, only those are read from disk.
Routed expert One of the many feed-forward experts a router can choose. Streamed from SSD rather than kept resident.
Shared expert An expert every token uses. It is always needed, so it stays resident with the dense backbone.
Top-k routing The router picks the k highest-scoring experts per token per layer. The set of indices it picks is a discrete decision, not a numeric value — see invariant 2 on the Architecture page.
Dense backbone Everything that is not a routed expert: attention, the router, the shared expert and the head. Replicated on every node under expert parallelism.
Dense model A model with no routed experts at all. M0 uses one deliberately, so the harness can be proven without routing in the way.
Sparse / compressed attention Attention variants that attend to a selected subset of positions or blocks instead of the whole context. The selection is discrete and must match the reference exactly.
Constant-size recurrent state A state that does not grow with context length, used by newer hybrid attention designs.
MTP (multi-token prediction) A speculative decoding scheme that drafts several tokens ahead for the main model to verify.
N-gram lookup table A large precomputed table used by some models to shortcut predictable spans. It is big, and how it is shared across nodes is an open question.
FP4 / FP8 / bf16 Weight and activation number formats. Frontier MoE models ship experts in FP4 or FP8 natively; bf16 is the format M0 works in.
Quantization Packing weights into fewer bits. This project prefers the vendor's own quantized weights over re-quantizing.
Transcoding Converting a quantized representation into the runtime's layout without dequantizing and re-quantizing. Invariant 4.
KV cache The stored keys and values of previous tokens. It grows with context and competes for the same memory as the expert cache.

Distributed terms

Term Meaning
Expert parallelism Sharding by expert: every node holds a disjoint 1/N slice of the routed experts and all nodes work on the same token. This project's core choice.
Pipeline parallelism Sharding by layer: each node owns a contiguous block of layers. Rejected here, because with one sequence in flight only one node is busy per token.
All-reduce Combining every node's partial MoE output into the result all nodes agree on. One per MoE layer, roughly 4 KB of payload, so latency dominates bandwidth.
Reduction order The order in which partial sums are combined. Floating-point addition is not associative, so the order must be fixed for results to be reproducible at all.
Bit-identical Byte-for-byte equal output. The M2 gate: two nodes must produce exactly what one node produced.
Discrete decision A choice from a set — a top-k index set, a sparse-attention block selection — as opposed to a numeric tensor. Checked for exact equality, separately from numeric tolerances.
Golden trace A recorded reference run: per-layer tensors, logits and discrete index sets, captured so a runtime's output can be diffed against it byte for byte.
Cold start First token of a run, before any expert is cached.

Runtime terms

Term Meaning
IR (intermediate representation) The declarative model description the runtime runs. It dispatches on tensor role, not tensor name, so architecture-specific knowledge lives only in importers.
Importer The per-family module that maps a checkpoint's tensor names onto IR roles. Intended to be a pure mapping, roughly 500–800 lines.
Transform pass An architecture-agnostic rewrite over the IR: fusing, repacking, expert reordering, quantization, sharding.
Shard policy The data that says which node owns which experts, how the dense backbone is replicated, and which policy version is in use.
Bounded expert cache The fixed memory budget an expert must fit in. Cache misses are SSD reads; the budget, not the page cache, is the working set.
SSD streaming Reading expert weights from local storage at the moment they are selected, instead of holding the model in memory.

Project terms

Term Meaning
ttd TinyTitan Datacenter: this project. Short name ttd. A dedicated repository, built from scratch and licensed under MIT. It is a the cluster.
TinyTitan Datacenter This project — the distributed, expert-parallel engine.
Milestone (M0–M5) A phase of the plan with a single gate. See the Roadmap.
Gate The measurement that closes a milestone. A phase is not done because the code exists; it is done when its gate passes on real hardware.
DC-nnn A stable task identifier in the Project Tracker. Never reused, never renumbered.
ADR Architecture decision record: the choice, the alternatives and why they were rejected.
TinyTitan Datacenter

Project

Design

Development

Sister project: TinyTitan

Clone this wiki locally