Repository navigation
Home
Important
Status: builds and passes its tests; M0, M1 and M2 are complete and gated; M3's
functional half is demonstrated and its throughput gate needs a quiet farm. On the
toolchain Package.swift requires — Swift 6.4 / Xcode 27 — the build is clean and
swift test --no-parallel runs 266 tests, 0 skipped, 0 failures, with 427
standard-library Python tests gated in CI. Xcode 27 with Swift 6.4 is the only
supported toolchain, and the Swift CI job refuses any other rather than warning and
skipping. M0 is complete and its gate passed: Qwen/Qwen3.5-2B, three frozen
prompts, 40,683,520 bytes identical to the contract with every discrete decision
matching the reference
(M0 gate).
M1's gate passes in both of its forms on all five frozen prompts — the engine against
a contract reading the same checkpoint (b8c976c5e7ba8816…) and against one reading the
same install (b0d382dbabf36df0…), every comparison 83 tensors, 0 differing elements and
40 discrete decisions. M2's gate passes across two machines on the real 35B model.
M4 and M5 have not started. v1.0.0 is released — arm64 binaries for macOS 26 or newer, built by
tools/release.py, with VERSION as the single source of the version and the
changelog as its announcement. Where this wiki and
docs/ disagree about a measurement, docs/ is right, and the
Project Tracker carries the task state.
TinyTitan Datacenter is a distributed inference engine for large mixture-of-experts (MoE) language models on a cluster of Mac minis and Mac Studios over LAN/SFP/QSFP and Thunderbolt. Expert weights are streamed from SSD, the model is split across the nodes, and the cluster is optimised for one user at a time rather than for serving throughput.
-
M0 is complete. Its gate passed on
Qwen/Qwen3.5-2B: three frozen prompts, 40,683,520 bytes of trace data identical to the contract, and every discrete decision matching the reference implementation. -
The toolchain is enforced, with no exceptions. Xcode 27 with Swift 6.4 is the only
supported pairing; CI fails on a runner that cannot provide it rather than warning and
skipping — and writing that check found a
tail -1bug that had made the old one skip on every runner it ever saw, including a correct one. -
I6 provenance is closed. Every install now records a sha256 per source weight file and
a commit or
null, with tests the previous behaviour cannot satisfy. -
M1's gate passes, in both of its forms. The checkpoint pair reproduces the digest the
2026-09-16 run recorded (
b8c976c5e7ba8816…) and the install pair — M1's restated claim — producesb0d382dbabf36df0…, both on all five frozen prompts. The cluster's tok/s measurement is what remains, and it needs a quiet farm.
Everything else, dated and with its evidence: News.
Splitting layers across machines does not help: with one sequence in flight only one node is ever busy, so bytes read per token is unchanged. This project shards by expert parallelism instead — every node holds the dense backbone replicated and a disjoint 1/N slice of the routed experts, all nodes work on the same token at the same time, and the MoE output is all-reduced.
| Pipeline parallel | Expert parallel | |
|---|---|---|
| Nodes busy per token | 1 of N | N of N |
| Aggregate SSD bandwidth | 1x | N x |
| Aggregate expert cache | 1x | N x |
| Sync per token | N-1 hops | 1 all-reduce per MoE layer (~4 KB) |
That is the whole thesis: 4–10x single-user tok/s from topology and I/O layout, not from more compute.
| Goal | Page |
|---|---|
| Read what has closed, what it cost and what it taught | News |
| See the plan, its phases and its gates | Roadmap |
| See what is planned, in progress, or blocked | Project Tracker |
| Read the technical design and its invariants | Architecture |
| See the verified configuration of the three target models | Target models |
| Check the hardware and toolchain the work assumes | Testbed |
| Look up a term used on these pages | Glossary |
| Milestone | What it is | Gate |
|---|---|---|
| M0 | Single node, small dense model, bf16 | Bit-matches reference golden traces |
| M1 | Qwen3.6-35B-A3B, single node, 4-bit, SSD-streamed | Correct output, recorded tok/s baseline |
| M2 | 2 nodes, expert-parallel | Bit-identical to M1 — the project's real gate |
| M3 | 4 nodes | ≥3x the M1 tok/s |
| M4 | DeepSeek-V4.1-Flash | Matches the reference at 128K context |
| M5 | Qwen3.8-Flash-Next | Same gate as M4 |
Details, dependencies and exit evidence: Roadmap · live status: Project Tracker.
ttd is a dedicated project, built from scratch and licensed under MIT. It is a native Swift and Metal runtime that runs Qwen-family MoE and dense models on Apple Silicon, streaming routed experts from SSD so a model larger than RAM still runs, measured at 7.1–7.9 tok/s decode on each of four Mac minis and 6.0 tok/s through the four-stage cluster chain.
- 4x Mac mini M2 (8 GB each) over Thunderbolt is the target configuration.
- The engine is designed to scale to an unlimited number of nodes of the same model type, not just four.
- Apple Silicon Macs, macOS, Swift and Metal. macOS only, by design.
Training or fine-tuning · multi-user serving and continuous batching · architectural modification of imported models · being a universal model translator. Each new attention family costs real kernel work, and that is expected.
| Source repository | Pummelchen/TinyTitan_Datacenter |
| Licence | MIT — see LICENSE |
| Contact | André Borchert — 0xa0b1@gmail.com |
Source · Issues · Discussions · Sister project: TinyTitan
Project
Design
Development
Sister project: TinyTitan