Repository navigation
Releases: Pummelchen/TinyTitan_Datacenter
Release list
TinyTitan Datacenter 1.1
TinyTitan Datacenter 1.1
ttd 1.1 is the first release under this project's own identity, and the number is its own.
TinyTitan Datacenter is a dedicated repository - not a fork, with no upstream - built from scratch
and licensed under MIT, and its release line starts from v1.0.0 rather than from any number
inherited with the code. The version is 1.1, two components: there is no patch component in this
project.
What is in it
The repository ships one runtime: a native Swift and Metal engine that streams routed experts from
SSD so a model larger than RAM still runs, plus the cluster layer that runs it across machines.
Single node, measured. Four Mac minis, one binary (md5 ec6710cd... on all four), 128 tokens at
temperature 0, prompt "The capital of France is":
| node | tok/s | | --- | --- | | node1 | 7.935 | | node2 | 7.937 | | node3 | 7.884 | | node4 | 7.146
/ 7.448 / 7.396 |
Mean 7.73, and every machine produces the same text.
Four nodes, measured. A four-stage layer pipeline runs one model across all four machines and
produces the reference text at 6.016 tok/s at 128 tokens. It is correct and balanced within 9% -
and it is slower than a single node, which this release states rather than dresses up.
What this release does not claim
21 tok/s is not reachable on this hardware as configured, and the reason is measured rather than
asserted:
- the step is 141.6 ms/token: 47% expert I/O and 50% GPU wait. Those two phases are
serial, not overlapped - the profiler's buckets sum to the step (66.5 + 71.3 + 3.8 = 141.6), and
taking 42.5 ms/token of expert-I/O await out of a run takes 42.8 ms of step with it, a
ratio of 1.007. Correction: an earlier draft read "the expert I/O is already fully
overlapped" on the strength ofexposed_ioreporting 0.0 ms on every node; that counter does
not answer whether the I/O adds to the step, and the arithmetic says it does. Found after v1.1
shipped, so the claim is withdrawn here; - a layer pipeline divides memory, not time: the head's period of 166.2 ms/token sits at
- the serial prediction (the sum of the stage times, 194.8 ms) and nowhere near the pipelined
- one (48.7 ms);
- every single-node tuning axis is at its optimum - prefetch depth, the prefetch ring, and the
- expert cache, where 40 slots is the best of seven points and 64 and above collapse into
- swap;
- the wire is 1 GbE: 118 MB/s and a 765 us round trip, and UDP measures the same as TCP, so
- it is the link and not the protocol. Pooling expert residency is **8.6x slower than local
- disk**, and tensor parallelism spends 61 ms of synchronisation against a 47.6 ms target.
The route to the target is a faster link, not more engine work. The Thunderbolt and 10GbE ports
on these machines reportstatus: inactive; connected, an RTT near 70 us puts the step near 26.6
ms.
Verification
- The four single-node runs and the four-stage chain above, each with its timing footer.
swift test --no-parallelon the runtime.- The release gates: force-cast ok, func-length ok (0 baselined, 0 new, 2186 scanned),
- unchecked-sendable ok, arch-path ok.
- A clean scratch release build. The archive is:
TinyTitan_Datacenter-1.1-macos-arm64.tar.gzsha256:
eb154f05b485c03a57d41f887d63f1906e7f447fb27dda1b5b6274a3a4864bfaTinyTitan_Datacenter-1.1-macos- arm64.tar.gzsize:25830004bytes - (Both fields are filled in at publish time. A clean rebuild is not byte-reproducible, so a
- digest or size quoted here would go stale the moment the archive was rebuilt.)
- No model, dataset or dependency was fetched to make anything pass.
Not checked, and named here as the gate requires
- Every golden baseline is NOT CHECKED. No install was present under
models/on the - machine that cut this release, and none may be fetched to change that:
ornith-8,ornith-4,qwen36-4,qwen36-8,qwen38-4,qwen38-8,agentworld-4,agentworld-8,katcoder-4,katcoder-8,qwen35-2b-4,qwen35-2b-8,qwen35-4b-4,qwen35-4b-8,qwen35-9b-4,qwen35-9b-8. **This release therefore ships no golden-baseline- evidence.** The benchmarks quoted above were measured through the CLI against an install
- outside
models/, which is a measurement, not a golden gate. - The converter-expert-order gate reports
SKIP: No module named 'numpy'— the converter's - dependencies are unavailable here, so that gate is not checked either.
TinyTitan Datacenter 1.0.0
TinyTitan Datacenter 1.0.0
Apple silicon (arm64) binaries for macOS 26 or newer.
[1.0.0] — 2026-09-17
Tag: v1.0.0
First release. The engine runs Qwen3.6-35B-A3B across four Apple-silicon nodes and produces the
single-node result exactly, with a measured decode speed-up of 1.7x over one node.
Highlights
- Distributed 35 B inference that is bit-identical to a single node. All four machines run one
forward in a full mesh and reproduce the single-node trace exactly: 83 tensors, 0 differing elements,
40 discrete decisions, matching digests (--mesh; M2's gate). - Vocabulary-parallel output head. The head was 1.05 s/step identical on every node — the largest
piece of replicated work. Each node now computes only its own vocabulary rows, inside the existing block
decomposition, and the slices are gathered so every node ends with the same full logits array: the
argmax, the margin and the trace digest are untouched.head1.045 → 0.26 s/step; cluster 1.13 → 1.36x
(D93;ShardedGenerateTestsasserts the gathered logits equal the single node's). - The int4 dequantiser runs across the machine's cores.
loadwas ~1 G parameters decoded with SIMD4
on one core of eight, on every token and every node. Rows are independent, so the row loop is spread;
SHARD_DECODE_THREADS=1restores the single-threaded path so the two are compared on one binary.
load1.33 → 0.56 s/step; one node 5.14 → 4.36 s/step (D94;Int4UnpackTestscompares the
vector path againstscalarbit-for-bit over more than fifty shapes). - Cluster decode at 1.70–1.74x over a single node, bit-identical on all four nodes, measured with the
conditions alternated on one binary. The farm was busy (loads 1.6–5.5) and load hurts the ratio, so these
are lower bounds (D94; the M3 gate, recorded as an observation — see Checks that did not run). - The exchange is measured in its parts. Encode, send, receive and merge are reported separately and
cover 100.0% of the exchange: receive is 99.6% of it, and the wire format and protocol cost 9 ms per
step. The imbalance hypothesis that followed was then falsified by four interleaved runs
(D92;ShardExchangeTests). - Shard plans can interleave expert ownership.
--distribution contiguous|round-robin, with unknown
values refused rather than defaulted (D92;test_run_m2_gate.py). - A version, single-sourced and enforced.
VERSIONis the authority,sources/DatacenterEngine/Version.swift
is generated from it,--versionanswers on all three tools, and bothtools/version.py --check(a gate)
andPackage.swift(at configure time) refuse a disagreement (RELEASE.md§1.3; 7 tests).
Requirements
- Apple silicon only (M1–M6). Built natively
arm64;lipo -archsreports exactlyarm64on every
binary in the archive. - macOS 26 or newer — the package's declared platform floor.
- A model install built by this repository's
tools/quantize.py, plus the plan file for a sharded run.
Binaries
datacenter-generate— generation and the throughput measurements.datacenter-trace— one forward, written as a trace.datacenter-node— one node of a sharded run, as its own process.- Not code-signed and not notarized. See
README-binaries.txtin the archive for the quarantine
command.
Checks that did not run
- M3's throughput gate was not asserted. It refuses a busy farm by design (
D38), and the farm was
shared throughout; every speed-up figure above is anobservation_onlyrun with its loads recorded
beside it. The gate's own roadmap target of ≥3x on a quiet farm is therefore not claimed. - The warning scan covers the release products, not the test targets. A clean release build of the
products is what this release scans; the test targets are built and run by the gate set in the debug
configuration, where a fresh scratch build passes all 226 tests. A release-configuration test build does
not resolveDatacenterIRon this toolchain, which is recorded as observed and not diagnosed. - No GPU matmul path is enabled.
MetalMatmulremains opt-in (SHARD_GPU_MATMUL=1) becauseD63
measured it slower than the CPU; the threadgroup-tiled kernel written since has not been re-measured,
so this release claims nothing about it.
Checksums
b6ebfd1b861a1ac434b50c87b43dda3e88513f760df7fd9942d8d886cdd72fdd TinyTitan_Datacenter-1.0.0-macos-arm64.tar.gz
2078299 bytes