Skip to content

Releases: Pummelchen/TinyTitan_Datacenter

TinyTitan Datacenter 1.1

Choose a tag to compare

@Pummelchen Pummelchen released this 19 Sep 23:38

TinyTitan Datacenter 1.1

ttd 1.1 is the first release under this project's own identity, and the number is its own.
TinyTitan Datacenter is a dedicated repository - not a fork, with no upstream - built from scratch
and licensed under MIT, and its release line starts from v1.0.0 rather than from any number
inherited with the code. The version is 1.1, two components: there is no patch component in this
project.

What is in it

The repository ships one runtime: a native Swift and Metal engine that streams routed experts from
SSD so a model larger than RAM still runs, plus the cluster layer that runs it across machines.
Single node, measured. Four Mac minis, one binary (md5 ec6710cd... on all four), 128 tokens at
temperature 0, prompt "The capital of France is":
| node | tok/s | | --- | --- | | node1 | 7.935 | | node2 | 7.937 | | node3 | 7.884 | | node4 | 7.146
/ 7.448 / 7.396 |
Mean 7.73, and every machine produces the same text.
Four nodes, measured. A four-stage layer pipeline runs one model across all four machines and
produces the reference text at 6.016 tok/s at 128 tokens. It is correct and balanced within 9% -
and it is slower than a single node, which this release states rather than dresses up.

What this release does not claim

21 tok/s is not reachable on this hardware as configured, and the reason is measured rather than
asserted:

  • the step is 141.6 ms/token: 47% expert I/O and 50% GPU wait. Those two phases are
    serial, not overlapped
    - the profiler's buckets sum to the step (66.5 + 71.3 + 3.8 = 141.6), and
    taking 42.5 ms/token of expert-I/O await out of a run takes 42.8 ms of step with it, a
    ratio of 1.007. Correction: an earlier draft read "the expert I/O is already fully
    overlapped" on the strength of exposed_io reporting 0.0 ms on every node; that counter does
    not answer whether the I/O adds to the step, and the arithmetic says it does. Found after v1.1
    shipped, so the claim is withdrawn here;
  • a layer pipeline divides memory, not time: the head's period of 166.2 ms/token sits at
  • the serial prediction (the sum of the stage times, 194.8 ms) and nowhere near the pipelined
  • one (48.7 ms);
  • every single-node tuning axis is at its optimum - prefetch depth, the prefetch ring, and the
  • expert cache, where 40 slots is the best of seven points and 64 and above collapse into
  • swap;
  • the wire is 1 GbE: 118 MB/s and a 765 us round trip, and UDP measures the same as TCP, so
  • it is the link and not the protocol. Pooling expert residency is **8.6x slower than local
  • disk**, and tensor parallelism spends 61 ms of synchronisation against a 47.6 ms target.
    The route to the target is a faster link, not more engine work. The Thunderbolt and 10GbE ports
    on these machines report status: inactive; connected, an RTT near 70 us puts the step near 26.6
    ms.

Verification

  • The four single-node runs and the four-stage chain above, each with its timing footer.
  • swift test --no-parallel on the runtime.
  • The release gates: force-cast ok, func-length ok (0 baselined, 0 new, 2186 scanned),
  • unchecked-sendable ok, arch-path ok.
  • A clean scratch release build. The archive is:
    TinyTitan_Datacenter-1.1-macos-arm64.tar.gz sha256:
    eb154f05b485c03a57d41f887d63f1906e7f447fb27dda1b5b6274a3a4864bfa TinyTitan_Datacenter-1.1-macos- arm64.tar.gz size: 25830004 bytes
  • (Both fields are filled in at publish time. A clean rebuild is not byte-reproducible, so a
  • digest or size quoted here would go stale the moment the archive was rebuilt.)
  • No model, dataset or dependency was fetched to make anything pass.

Not checked, and named here as the gate requires

  • Every golden baseline is NOT CHECKED. No install was present under models/ on the
  • machine that cut this release, and none may be fetched to change that:
  • ornith-8, ornith-4, qwen36-4, qwen36-8, qwen38-4, qwen38-8, agentworld-4,
  • agentworld-8, katcoder-4, katcoder-8, qwen35-2b-4, qwen35-2b-8, qwen35-4b-4,
  • qwen35-4b-8, qwen35-9b-4, qwen35-9b-8. **This release therefore ships no golden-baseline
  • evidence.** The benchmarks quoted above were measured through the CLI against an install
  • outside models/, which is a measurement, not a golden gate.
  • The converter-expert-order gate reports SKIP: No module named 'numpy' — the converter's
  • dependencies are unavailable here, so that gate is not checked either.

TinyTitan Datacenter 1.0.0

Choose a tag to compare

@Pummelchen Pummelchen released this 17 Sep 13:29

TinyTitan Datacenter 1.0.0

Apple silicon (arm64) binaries for macOS 26 or newer.

[1.0.0] — 2026-09-17

Tag: v1.0.0

First release. The engine runs Qwen3.6-35B-A3B across four Apple-silicon nodes and produces the
single-node result exactly, with a measured decode speed-up of 1.7x over one node.

Highlights

  • Distributed 35 B inference that is bit-identical to a single node. All four machines run one
    forward in a full mesh and reproduce the single-node trace exactly: 83 tensors, 0 differing elements,
    40 discrete decisions, matching digests (--mesh; M2's gate).
  • Vocabulary-parallel output head. The head was 1.05 s/step identical on every node — the largest
    piece of replicated work. Each node now computes only its own vocabulary rows, inside the existing block
    decomposition, and the slices are gathered so every node ends with the same full logits array: the
    argmax, the margin and the trace digest are untouched. head 1.045 → 0.26 s/step; cluster 1.13 → 1.36x
    (D93; ShardedGenerateTests asserts the gathered logits equal the single node's).
  • The int4 dequantiser runs across the machine's cores. load was ~1 G parameters decoded with SIMD4
    on one core of eight, on every token and every node. Rows are independent, so the row loop is spread;
    SHARD_DECODE_THREADS=1 restores the single-threaded path so the two are compared on one binary.
    load 1.33 → 0.56 s/step; one node 5.14 → 4.36 s/step (D94; Int4UnpackTests compares the
    vector path against scalar bit-for-bit over more than fifty shapes).
  • Cluster decode at 1.70–1.74x over a single node, bit-identical on all four nodes, measured with the
    conditions alternated on one binary. The farm was busy (loads 1.6–5.5) and load hurts the ratio, so these
    are lower bounds (D94; the M3 gate, recorded as an observation — see Checks that did not run).
  • The exchange is measured in its parts. Encode, send, receive and merge are reported separately and
    cover 100.0% of the exchange: receive is 99.6% of it, and the wire format and protocol cost 9 ms per
    step. The imbalance hypothesis that followed was then falsified by four interleaved runs
    (D92; ShardExchangeTests).
  • Shard plans can interleave expert ownership. --distribution contiguous|round-robin, with unknown
    values refused rather than defaulted (D92; test_run_m2_gate.py).
  • A version, single-sourced and enforced. VERSION is the authority, sources/DatacenterEngine/Version.swift
    is generated from it, --version answers on all three tools, and both tools/version.py --check (a gate)
    and Package.swift (at configure time) refuse a disagreement (RELEASE.md §1.3; 7 tests).

Requirements

  • Apple silicon only (M1–M6). Built natively arm64; lipo -archs reports exactly arm64 on every
    binary in the archive.
  • macOS 26 or newer — the package's declared platform floor.
  • A model install built by this repository's tools/quantize.py, plus the plan file for a sharded run.

Binaries

  • datacenter-generate — generation and the throughput measurements.
  • datacenter-trace — one forward, written as a trace.
  • datacenter-node — one node of a sharded run, as its own process.
  • Not code-signed and not notarized. See README-binaries.txt in the archive for the quarantine
    command.

Checks that did not run

  • M3's throughput gate was not asserted. It refuses a busy farm by design (D38), and the farm was
    shared throughout; every speed-up figure above is an observation_only run with its loads recorded
    beside it. The gate's own roadmap target of ≥3x on a quiet farm is therefore not claimed.
  • The warning scan covers the release products, not the test targets. A clean release build of the
    products is what this release scans; the test targets are built and run by the gate set in the debug
    configuration, where a fresh scratch build passes all 226 tests. A release-configuration test build does
    not resolve DatacenterIR on this toolchain, which is recorded as observed and not diagnosed.
  • No GPU matmul path is enabled. MetalMatmul remains opt-in (SHARD_GPU_MATMUL=1) because D63
    measured it slower than the CPU; the threadgroup-tiled kernel written since has not been re-measured,
    so this release claims nothing about it.

Checksums

b6ebfd1b861a1ac434b50c87b43dda3e88513f760df7fd9942d8d886cdd72fdd  TinyTitan_Datacenter-1.0.0-macos-arm64.tar.gz
2078299  bytes