Skip to content
Pummelchen edited this page Sep 19, 2026 · 31 revisions
TinyTitan Datacenter

TinyTitan Datacenter Wiki

Important

Status: builds and passes its tests; M0, M1 and M2 are complete and gated; M3's functional half is demonstrated and its throughput gate needs a quiet farm. On the toolchain Package.swift requires — Swift 6.4 / Xcode 27 — the build is clean and swift test --no-parallel runs 266 tests, 0 skipped, 0 failures, with 427 standard-library Python tests gated in CI. Xcode 27 with Swift 6.4 is the only supported toolchain, and the Swift CI job refuses any other rather than warning and skipping. M0 is complete and its gate passed: Qwen/Qwen3.5-2B, three frozen prompts, 40,683,520 bytes identical to the contract with every discrete decision matching the reference (M0 gate). M1's gate passes in both of its forms on all five frozen prompts — the engine against a contract reading the same checkpoint (b8c976c5e7ba8816…) and against one reading the same install (b0d382dbabf36df0…), every comparison 83 tensors, 0 differing elements and 40 discrete decisions. M2's gate passes across two machines on the real 35B model. M4 and M5 have not started. v1.0.0 is released — arm64 binaries for macOS 26 or newer, built by tools/release.py, with VERSION as the single source of the version and the changelog as its announcement. Where this wiki and docs/ disagree about a measurement, docs/ is right, and the Project Tracker carries the task state.

TinyTitan Datacenter is a distributed inference engine for large mixture-of-experts (MoE) language models on a cluster of Mac minis and Mac Studios over LAN/SFP/QSFP and Thunderbolt. Expert weights are streamed from SSD, the model is split across the nodes, and the cluster is optimised for one user at a time rather than for serving throughput.

Recent news

  • M0 is complete. Its gate passed on Qwen/Qwen3.5-2B: three frozen prompts, 40,683,520 bytes of trace data identical to the contract, and every discrete decision matching the reference implementation.
  • The toolchain is enforced, with no exceptions. Xcode 27 with Swift 6.4 is the only supported pairing; CI fails on a runner that cannot provide it rather than warning and skipping — and writing that check found a tail -1 bug that had made the old one skip on every runner it ever saw, including a correct one.
  • I6 provenance is closed. Every install now records a sha256 per source weight file and a commit or null, with tests the previous behaviour cannot satisfy.
  • M1's gate passes, in both of its forms. The checkpoint pair reproduces the digest the 2026-09-16 run recorded (b8c976c5e7ba8816…) and the install pair — M1's restated claim — produces b0d382dbabf36df0…, both on all five frozen prompts. The cluster's tok/s measurement is what remains, and it needs a quiet farm.

Everything else, dated and with its evidence: News.

The thesis

Splitting layers across machines does not help: with one sequence in flight only one node is ever busy, so bytes read per token is unchanged. This project shards by expert parallelism instead — every node holds the dense backbone replicated and a disjoint 1/N slice of the routed experts, all nodes work on the same token at the same time, and the MoE output is all-reduced.

Pipeline parallel Expert parallel
Nodes busy per token 1 of N N of N
Aggregate SSD bandwidth 1x N x
Aggregate expert cache 1x N x
Sync per token N-1 hops 1 all-reduce per MoE layer (~4 KB)

That is the whole thesis: 4–10x single-user tok/s from topology and I/O layout, not from more compute.

Where to go

Goal Page
Read what has closed, what it cost and what it taught News
See the plan, its phases and its gates Roadmap
See what is planned, in progress, or blocked Project Tracker
Read the technical design and its invariants Architecture
See the verified configuration of the three target models Target models
Check the hardware and toolchain the work assumes Testbed
Look up a term used on these pages Glossary

Milestones at a glance

Milestone What it is Gate
M0 Single node, small dense model, bf16 Bit-matches reference golden traces
M1 Qwen3.6-35B-A3B, single node, 4-bit, SSD-streamed Correct output, recorded tok/s baseline
M2 2 nodes, expert-parallel Bit-identical to M1 — the project's real gate
M3 4 nodes ≥3x the M1 tok/s
M4 DeepSeek-V4.1-Flash Matches the reference at 128K context
M5 Qwen3.8-Flash-Next Same gate as M4

Details, dependencies and exit evidence: Roadmap · live status: Project Tracker.

What it builds on

ttd is a dedicated project, built from scratch and licensed under MIT. It is a native Swift and Metal runtime that runs Qwen-family MoE and dense models on Apple Silicon, streaming routed experts from SSD so a model larger than RAM still runs, measured at 7.1–7.9 tok/s decode on each of four Mac minis and 6.0 tok/s through the four-stage cluster chain.

Development target

  • 4x Mac mini M2 (8 GB each) over Thunderbolt is the target configuration.
  • The engine is designed to scale to an unlimited number of nodes of the same model type, not just four.
  • Apple Silicon Macs, macOS, Swift and Metal. macOS only, by design.

Non-goals

Training or fine-tuning · multi-user serving and continuous batching · architectural modification of imported models · being a universal model translator. Each new attention family costs real kernel work, and that is expected.

Related

Source repository Pummelchen/TinyTitan_Datacenter
Licence MIT — see LICENSE
Contact André Borchert — 0xa0b1@gmail.com
TinyTitan Datacenter

Project

Design

Development

Sister project: TinyTitan

Clone this wiki locally