Skip to content

Architecture

Pummelchen edited this page Sep 19, 2026 · 5 revisions
TinyTitan Datacenter

Architecture

The design as it stands. Nothing on this page is built: the repository is in design phase. Where a mechanism is not yet decided, it says so and points at the open question in the Project Tracker rather than inventing an answer.

Shape of the system

                    ┌───────────── coordinator ─────────────┐
prompt ──▶ token ──▶ │ dense backbone (replicated per node)  │
                    │   attention · router · shared expert  │
                    └───┬─────────────┬─────────────┬───────┘
                        │             │             │
                   ┌────▼───┐    ┌────▼───┐    ┌────▼───┐
                   │ node 1 │    │ node 2 │ …  │ node N │   experts 1/N each,
                   │ E-slice│    │ E-slice│    │ E-slice│   streamed from SSD
                   └────┬───┘    └────┬───┘    └────┬───┘
                        └──────── all-reduce ───────┘
                              (fp32, fixed order)
                                      │
                                      ▼
                                next token

Sharding model: expert parallelism

Every node holds the dense backbone replicated — attention, the router and the shared expert — plus a disjoint 1/N slice of the routed experts. All nodes process the same token simultaneously; each computes the contribution of the experts it owns; the MoE output is combined with one all-reduce per MoE layer.

The alternative, pipeline parallelism, is rejected because with a single sequence in flight only one node is ever busy: bytes read per token does not change, so neither does throughput. See the comparison table on Home.

Consequences that follow from this choice:

  • The dense backbone is duplicated N times. Its size is therefore a per-node constraint, not a cluster one, on 8 GB machines in particular.
  • Aggregate expert cache and aggregate SSD bandwidth scale with N, which is where the speedup comes from.
  • Routing statistics decide whether the split is balanced; a slice that owns few popular experts is idle more often. This is deliberately validated on measured routing, not assumed.

Invariants

These are the properties the design is built around. Each one is a gate, not an aspiration.

1. Bit-reproducibility is a hard invariant. N-node output must be bit-identical to 1-node output. That means fp32 accumulation and a fixed reduction order — the all-reduce cannot be allowed to sum in a different order on different runs. Without this, a conversion bug and a scheduling artifact are indistinguishable, and on these models you will need to tell them apart. See DC-005 and DC-006.

2. Discrete decisions are checked separately from numerics. Router top-k index sets and sparse-attention block selections must match the reference exactly. Small numeric drift flips them, after which the output diverges completely while every per-tensor MSE check still looks green. A gate that only measures smoothness is a gate that cannot see the failure mode this project is most exposed to.

3. Faithful ports only — no architectural changes. No fewer layers, no weight sharing, no substituted attention. All speed comes from sharding, expert repacking and quantization. This keeps a reference implementation to diff against, which is what makes the project debuggable at all.

4. Transcode, don't requantize. DeepSeek ships FP4 experts and FP8 dense weights natively, and Qwen ships an FP8 variant. The path converts those representations directly; dequantizing and re-quantizing would stack this project's error on top of the vendor's.

One IR, thin importers

A declarative model IR dispatches on tensor role, not tensor name. The pieces:

Piece Responsibility Cost
Importer pure name-to-role mapping for one model family ~500–800 lines per family
Transform passes fusing, repacking, expert reordering, quantization, sharding written once, architecture-agnostic
Quantization and shard policy data files shared across models data, not code

The intended payoff: a new model in an existing family is a conversion job, and a new family costs an importer plus whatever new attention kernel work it genuinely needs. An IR built around the current frontier shape — fine-grained MoE, shared expert, compressed or sparse hybrid attention, constant-size recurrent state, MTP speculative decoding, sparse n-gram lookup tables — should absorb the next generation with importer changes only.

Whether this IR is shared with the sister runtime or built beside it is not yet decided (Q2, DC-010).

Streaming and cache

Routed expert weights are not resident: they are read from SSD when selected, into a bounded per-layer cache. The cache budget is the actual working set, so it is a tuning parameter and a memory-safety boundary rather than a hint. The runtime already demonstrates this pattern single-node, including overlapping cache-hit execution with in-flight miss reads; this project's question is what changes when misses are spread across N machines with one SSD each.

Related: the M3 phase measures the synchronisation budget, because the win comes from reads that are hidden by work, not from reads that are merely parallel (DC-051).

Transport and synchronisation

  • One all-reduce per MoE layer, carrying roughly 4 KB — the payload is small, so latency per layer, not bandwidth, is the number that matters.
  • Raw sockets over the plan's transports: LAN, SFP/QSFP, or the Thunderbolt bridge.
  • Development currently runs over 1 Gbit Ethernet — a developer limitation, not the design target — so early cluster timing is a floor rather than what the documented transports would give (R4, DC-008).
  • Which transport is mandatory, and what macOS permits over a bridge, is not yet decided and will be settled by measurement (Q5, DC-008).
  • Failure semantics for a single-user run — what a user sees when one node of four dies — are part of the M2 phase (DC-043).

Stack and platform

Language and kernels Swift 6.4 + Metal on Xcode 27, hand-written kernels (no MLX dependency in the sister runtime)
Platform macOS on Apple Silicon only, by design
Model format this project's own verified install format and repacker
Cluster scale designed for an unlimited number of nodes of the same model type, developed on four

What this page is not

There is no measured number on this page, and there is no code behind it. When a claim here is tested, the result belongs in the Project Tracker with its evidence — including the results that contradict the design.

Security posture

Stated here so it is a decision rather than an assumption. DC-083 audited it, D30 records what the audit found.

What is exposed. A node binds one address and one port, taken from its own cluster config, and connects only to peers that config names — the mesh rule is connect-to-lower, accept-from-higher, and nothing else listens. A node is not a listener to the world, and the addresses it uses are the private segment's. Model installs and the cluster config are read from the local filesystem; no part of the protocol fetches anything.

What the wire refuses, before it trusts anything: a frame whose magic or version is wrong, a dimension or term count outside its bound, a term whose width disagrees with its own header, a token outside the declared range, trailing bytes, a frame longer than 64 MB — and, from D30, a frame whose declared geometry disagrees with the receiver's, which is a check only the receiver can make. The reduction refuses a duplicate key with different bits, and refuses to sum an incomplete set of a token's selected experts.

What it does not do. There is no authentication and no encryption: any process that can reach the port can speak the protocol. That is deliberate — this targets a switched, private segment, and the operational rule is that it runs there — but it is a boundary of the design rather than a property of it. Running the same protocol anywhere less trusted would need mutual authentication and a keyed transcript first, and this section is where that would be recorded.

TinyTitan Datacenter

Project

Design

Development

Sister project: TinyTitan

Clone this wiki locally