Skip to content

Testbed

Pummelchen edited this page Sep 19, 2026 · 11 revisions
TinyTitan Datacenter

Testbed

What the work is developed and verified on, and how a node is checked before a run. Nothing here has been built or benchmarked yet — this page describes the target configuration and the readiness checks, not measured results.

Important

No operational secrets belong on this page. Host names, addresses, passwords, keys and model-access tokens are deliberately absent — this wiki is public and stays that way (R9, DC-083). Nodes are referred to by role: node 1 … node N.

Target configurations

Configuration Purpose Notes
1 Mac (developer machine) M0 and M1 — single node, harness, throughput baseline Same hardware class as a cluster node wherever possible, so the baseline means something (R8)
2 nodes M2 — the bit-identity gate The smallest configuration that can prove sharding is transparent
4x Mac mini M2, 8 GB each M3 — the development target The configuration the design is written against
N nodes M4/M5 and beyond Same model type per cluster; the engine is intended to scale past four

Class of hardware assumed: Apple Silicon Mac minis and Mac Studios, each with its own SSD holding that node's expert shard, connected over LAN, SFP/QSFP or a Thunderbolt bridge. Which of those is mandatory is an open decision (Q5, DC-008). The farm as inventoried on 2026-09-15 is four identical Mac mini M2 machines with 8 GB each on a switched 1 Gbit segment — see Farm inventory below.

Note

The link actually in use during development is 1 Gbit Ethernet. That is a developer limitation for now, not the design target: the transports above are what the design assumes and what DC-008 measures. Early cluster timing is a floor for that reason — and because the all-reduce payload is small (roughly 4 KB per MoE layer), the link shows up as latency before it shows up as bandwidth. The measured latencies of both available paths are in the inventory.

What a node must provide

Requirement Why
macOS on Apple Silicon, same major version across the cluster Kernel and Metal behaviour must not differ per node
Swift 6.4 toolchain and Xcode 27 with the Metal toolchain Kernels are compiled per node, so the toolchain version is part of the node contract
The same build of this project on every node A version hash mismatch must refuse to start, not mis-compute (DC-042)
The same model install, verified on the node An install receipt is bound to its absolute path; a moved install fails to load
Free SSD space for the node's shard Experts stream from local storage; this is the resource the design leans on
Reachability to every peer with the chosen transport Measured, not assumed (DC-008)
8 GB of unified memory (target class) The dense backbone, KV state, expert cache and macOS all fit inside it — the binding constraint (R3)

Toolchain, stated once so a node check has something to compare against: Swift 6.4 (the reference machine reports swiftlang-6.4.0.34.1, 2026-09-15), Xcode 27.0 with the Metal toolchain, and swift-tools-version:6.4 in the manifest once it exists. The per-feature register — which upcoming 6.4 features are enabled and which are deliberately not — is DC-015.

Python: CPython 3.14 is the project's Python standard and what the farm runs (see the inventory below). The gates under tools/ are standard-library-only, so they need no packages at all; Core ML work uses a project venv (.venv, CPython 3.14) built from tools/requirements-coreml.txt.

coremltools there is pinned to 9.1.dev1, a pre-release, and pinned exactly because of that: no stable release has a CPython 3.14 wheel yet — 9.0 covers CPython 3.10–3.13, and 9.1.dev1 (published 2026-08-03) is the only cp314 arm64 wheel. A conversion toolchain that changed under the project would make a converted model unattributable to the version that produced it. When a stable 9.1 lands, the pin moves to it.

Verified on the reference machine, 2026-09-15, on CPython 3.14.7: a MIL program converts, saves as an .mlpackage, loads, and predicts [3.0, 4.0, 5.0, 6.0] on both CPU_ONLY and CPU_AND_NE. Whether the Neural Engine belongs in a node's engine is still open (Q8); installing the toolchain does not decide it.

Farm inventory

Measured on 2026-09-15, from this checkout's machine, over SSH to every node. All four nodes are identical, which is the point of the exercise: a difference in memory, macOS build or toolchain would have to be found now, not in the middle of an M2 gate.

Node Internal LAN Model Chip Memory macOS Swift / Xcode Python
node1 192.168.18.27 Mac14,3 Apple M2 8 GB 27.0 (26A428) 6.4 / 27.0 3.14.7
node2 192.168.18.25 Mac14,3 Apple M2 8 GB 27.0 (26A428) 6.4 / 27.0 3.14.7
node3 192.168.18.29 Mac14,3 Apple M2 8 GB 27.0 (26A428) 6.4 / 27.0 3.14.7
node4 192.168.18.26 Mac14,3 Apple M2 8 GB 27.0 (26A428) 6.4 / 27.0 3.14.7

Cluster runs use these internal addresses, not the node names. The names resolve over the mesh VPN, which measures ~2.5x the round trip of the segment and is the slow way round the trap below; the addresses above are the switched Ethernet segment. An operator gave permission to record them here, so the next run does not have to rediscover them.

Node 4 is the machine this checkout lives on. Every node carries 228 GB of SSD, two Thunderbolt ports, Homebrew and uv; none of them has coremltools installed — that lives only in this repository's venv, because whether a node's engine needs Core ML at all is still Q8.

Standing operational note

The nodes are shared with other work, so the standing rule for cluster runs in the development phase is functional tests rather than benchmarks: a bit-identity check over the fixture moves a few megabytes and finishes in milliseconds, while a sustained throughput run competes with whatever else the machine is doing. Benchmarks are taken deliberately, on a quiet farm, and the numbers say so when they are taken.

For a real-model cluster run the model install has to exist on the node serving that shard. For M2's gate it was staged once, into a node's Downloads folder, because the development host already had it — 21.7 GB over the LAN, the only data movement a two-node gate needs. A node reads only the tensors its shard touches, so a full install per node is more than the design requires: a repacked per-shard install is a later refinement, not a prerequisite.

Network, as measured

  • Each node has two paths to the others: a switched 1 Gbit Ethernet segment (en0, 1000baseT <full-duplex>) and a mesh VPN, over which the node names resolve. The names take the VPN path — following the name is therefore the slow way, and this is the trap to remember: a cluster run that binds whatever a host name resolves to silently gets the VPN.
  • Round-trip latency, 5 packets each, warm: 0.49–0.64 ms over the Ethernet segment versus 1.4–1.8 ms over the VPN, measured from node 4; node 1 to node 2 measured 0.59 ms on the segment. Both are usable; neither is equal, and the difference is ~2.5x on the one number the all-reduce actually spends.
  • No Thunderbolt bridge is configured on any node (bridge0 is down), although each node has two Thunderbolt ports. The current development link is the 1 Gbit segment, as stated above it.
  • Addresses, credentials and access paths are deliberately not recorded here, and neither is the VPN's provider configuration (R9). The measurements above are the part that matters for the design.

Measured baselines, 2026-09-15

Taken on this checkout's machine and on node 1, with probes that assert what they claim (a dead echo server, a cached file and a wrong path all produce plausible-looking numbers, so each probe below was built to fail rather than to flatter):

Quantity Method Measured
ICMP round trip, 1 Gbit segment 20 pings at 0.2 s min 0.285 / avg 0.515 / max 0.990 ms, stddev 0.136
4 KB ping-pong, TCP, segment 200 exchanges, warm, TCP_NODELAY, identity checked p50 645 µs, p99 1157 µs
4 KB ping-pong, UDP, segment 200 exchanges, warm p50 687 µs, p99 1370 µs
4 KB ping-pong, loopback control same code path, 127.0.0.1 p50 25 µs
Internal SSD, sequential read 10.7 GB, 1 MiB blocks, F_NOCACHE on write and read, iostat confirming driver traffic 1161 MB/s (driver peaks 1574 MB/s)
Internal SSD, random 16 KB read, QD1 2000 aligned random reads, F_NOCACHE 6618 IOPS = 108 MB/s
Internal SSD, sequential write 10.7 GB, 1 MiB blocks, F_NOCACHE 641 MB/s
Memory immediately reclaimable, idle 8 GB node vm_stat free + inactive + speculative ~2.5 GB (macOS reports 63% "free")
External NVMe per node diskutil list external none attached — the 3 GB/s assumption has no hardware behind it yet

The two numbers that change decisions: the transport is base-RTT-bound, not bandwidth-bound (4 KB costs 645 µs against a 515 µs empty round trip, so the payload is nearly free and UDP buys nothing), and random 16 KB reads at queue depth 1 are 9–11x slower than sequential reads (108 vs 1161 MB/s), which is the entire case for a deep read queue and expert caching. The derived per-token I/O and synchronisation costs are worked out on the Target models page.

Node readiness checklist

Run before a cluster run, on each node. These are read-only checks; expected values are recorded in the Project Tracker as evidence, not in this wiki.

sw_vers                      # macOS version — identical across nodes
swift --version              # toolchain — identical across nodes
xcodebuild -version          # Xcode / Metal toolchain
python3 --version            # 3.14 across the farm
sysctl -n hw.memsize         # unified memory
sysctl -n machdep.cpu.brand_string
df -h /                      # SSD headroom for this node's shard
ping -c 5 <peer-name>        # the name takes the VPN path — measured 1.4–1.8 ms
ping -c 5 <peer-address>     # the Ethernet segment — measured 0.49–0.64 ms

Reachability is checked twice on purpose. On this farm the node names resolve over the VPN while the direct segment is reached by address, so a check that only follows names measures the wrong path — and gets a number 2.5x too pessimistic for the all-reduce (R4).

Also verify, once the code exists: identical build version hash, identical model install digest, and that the node can read its own shard at the expected rate.

Cross-checks and reference runs

A second, independent implementation is needed to produce golden traces — bit-exactness has no meaning without a reference (R7, DC-021). Options, none chosen yet:

  • a CPU reference path,
  • a separate reference implementation run on the developer machine,
  • a virtualised or remote reference runner for an independent result.

Virtualisation is available on the farm and on the developer machine (a hypervisor on one Mac, and QEMU elsewhere) if an independent platform is wanted for a cross-check. A virtualised reference must never be the only evidence for a gate: the gate is passed on real hardware.

TinyTitan Datacenter

Project

Design

Development

Sister project: TinyTitan

Clone this wiki locally