Repository navigation
Testbed
What the work is developed and verified on, and how a node is checked before a run. Nothing here has been built or benchmarked yet — this page describes the target configuration and the readiness checks, not measured results.
Important
No operational secrets belong on this page. Host names, addresses, passwords, keys
and model-access tokens are deliberately absent — this wiki is public and stays that
way (R9, DC-083). Nodes are referred to by role: node 1 … node N.
| Configuration | Purpose | Notes |
|---|---|---|
| 1 Mac (developer machine) | M0 and M1 — single node, harness, throughput baseline | Same hardware class as a cluster node wherever possible, so the baseline means something (R8) |
| 2 nodes | M2 — the bit-identity gate | The smallest configuration that can prove sharding is transparent |
| 4x Mac mini M2, 8 GB each | M3 — the development target | The configuration the design is written against |
| N nodes | M4/M5 and beyond | Same model type per cluster; the engine is intended to scale past four |
Class of hardware assumed: Apple Silicon Mac minis and Mac Studios, each with its own SSD holding that node's expert shard, connected over LAN, SFP/QSFP or a Thunderbolt bridge. Which of those is mandatory is an open decision (Q5, DC-008). The farm as inventoried on 2026-09-15 is four identical Mac mini M2 machines with 8 GB each on a switched 1 Gbit segment — see Farm inventory below.
Note
The link actually in use during development is 1 Gbit Ethernet. That is a developer limitation for now, not the design target: the transports above are what the design assumes and what DC-008 measures. Early cluster timing is a floor for that reason — and because the all-reduce payload is small (roughly 4 KB per MoE layer), the link shows up as latency before it shows up as bandwidth. The measured latencies of both available paths are in the inventory.
| Requirement | Why |
|---|---|
| macOS on Apple Silicon, same major version across the cluster | Kernel and Metal behaviour must not differ per node |
| Swift 6.4 toolchain and Xcode 27 with the Metal toolchain | Kernels are compiled per node, so the toolchain version is part of the node contract |
| The same build of this project on every node | A version hash mismatch must refuse to start, not mis-compute (DC-042) |
| The same model install, verified on the node | An install receipt is bound to its absolute path; a moved install fails to load |
| Free SSD space for the node's shard | Experts stream from local storage; this is the resource the design leans on |
| Reachability to every peer with the chosen transport | Measured, not assumed (DC-008) |
| 8 GB of unified memory (target class) | The dense backbone, KV state, expert cache and macOS all fit inside it — the binding constraint (R3) |
Toolchain, stated once so a node check has something to compare against:
Swift 6.4 (the reference machine reports swiftlang-6.4.0.34.1, 2026-09-15),
Xcode 27.0 with the Metal toolchain, and swift-tools-version:6.4 in the manifest
once it exists. The per-feature register — which upcoming 6.4 features are enabled and
which are deliberately not — is DC-015.
Python: CPython 3.14 is the project's Python standard and what the farm runs
(see the inventory below). The gates under tools/ are standard-library-only, so they
need no packages at all; Core ML work uses a project venv (.venv, CPython 3.14) built
from tools/requirements-coreml.txt.
coremltools there is pinned to 9.1.dev1, a pre-release, and pinned exactly
because of that: no stable release has a CPython 3.14 wheel yet — 9.0 covers CPython
3.10–3.13, and 9.1.dev1 (published 2026-08-03) is the only cp314 arm64 wheel. A
conversion toolchain that changed under the project would make a converted model
unattributable to the version that produced it. When a stable 9.1 lands, the pin moves
to it.
Verified on the reference machine, 2026-09-15, on CPython 3.14.7: a MIL program
converts, saves as an .mlpackage, loads, and predicts [3.0, 4.0, 5.0, 6.0] on both
CPU_ONLY and CPU_AND_NE. Whether the Neural Engine belongs in a node's engine is
still open (Q8); installing the toolchain does not decide it.
Measured on 2026-09-15, from this checkout's machine, over SSH to every node. All four nodes are identical, which is the point of the exercise: a difference in memory, macOS build or toolchain would have to be found now, not in the middle of an M2 gate.
| Node | Internal LAN | Model | Chip | Memory | macOS | Swift / Xcode | Python |
|---|---|---|---|---|---|---|---|
| node1 | 192.168.18.27 |
Mac14,3 | Apple M2 | 8 GB | 27.0 (26A428) | 6.4 / 27.0 | 3.14.7 |
| node2 | 192.168.18.25 |
Mac14,3 | Apple M2 | 8 GB | 27.0 (26A428) | 6.4 / 27.0 | 3.14.7 |
| node3 | 192.168.18.29 |
Mac14,3 | Apple M2 | 8 GB | 27.0 (26A428) | 6.4 / 27.0 | 3.14.7 |
| node4 | 192.168.18.26 |
Mac14,3 | Apple M2 | 8 GB | 27.0 (26A428) | 6.4 / 27.0 | 3.14.7 |
Cluster runs use these internal addresses, not the node names. The names resolve over the mesh VPN, which measures ~2.5x the round trip of the segment and is the slow way round the trap below; the addresses above are the switched Ethernet segment. An operator gave permission to record them here, so the next run does not have to rediscover them.
Node 4 is the machine this checkout lives on. Every node carries 228 GB of SSD, two
Thunderbolt ports, Homebrew and uv; none of them has coremltools installed — that
lives only in this repository's venv, because whether a node's engine needs Core ML at
all is still Q8.
The nodes are shared with other work, so the standing rule for cluster runs in the development phase is functional tests rather than benchmarks: a bit-identity check over the fixture moves a few megabytes and finishes in milliseconds, while a sustained throughput run competes with whatever else the machine is doing. Benchmarks are taken deliberately, on a quiet farm, and the numbers say so when they are taken.
For a real-model cluster run the model install has to exist on the node serving that shard. For M2's
gate it was staged once, into a node's Downloads folder, because the development host already had it —
21.7 GB over the LAN, the only data movement a two-node gate needs. A node reads only the tensors its
shard touches, so a full install per node is more than the design requires: a repacked per-shard install
is a later refinement, not a prerequisite.
- Each node has two paths to the others: a switched 1 Gbit Ethernet segment
(
en0,1000baseT <full-duplex>) and a mesh VPN, over which the node names resolve. The names take the VPN path — following the name is therefore the slow way, and this is the trap to remember: a cluster run that binds whatever a host name resolves to silently gets the VPN. - Round-trip latency, 5 packets each, warm: 0.49–0.64 ms over the Ethernet segment versus 1.4–1.8 ms over the VPN, measured from node 4; node 1 to node 2 measured 0.59 ms on the segment. Both are usable; neither is equal, and the difference is ~2.5x on the one number the all-reduce actually spends.
-
No Thunderbolt bridge is configured on any node (
bridge0is down), although each node has two Thunderbolt ports. The current development link is the 1 Gbit segment, as stated above it. - Addresses, credentials and access paths are deliberately not recorded here, and neither is the VPN's provider configuration (R9). The measurements above are the part that matters for the design.
Taken on this checkout's machine and on node 1, with probes that assert what they claim (a dead echo server, a cached file and a wrong path all produce plausible-looking numbers, so each probe below was built to fail rather than to flatter):
| Quantity | Method | Measured |
|---|---|---|
| ICMP round trip, 1 Gbit segment | 20 pings at 0.2 s | min 0.285 / avg 0.515 / max 0.990 ms, stddev 0.136 |
| 4 KB ping-pong, TCP, segment | 200 exchanges, warm, TCP_NODELAY, identity checked |
p50 645 µs, p99 1157 µs |
| 4 KB ping-pong, UDP, segment | 200 exchanges, warm | p50 687 µs, p99 1370 µs |
| 4 KB ping-pong, loopback control | same code path, 127.0.0.1
|
p50 25 µs |
| Internal SSD, sequential read | 10.7 GB, 1 MiB blocks, F_NOCACHE on write and read, iostat confirming driver traffic |
1161 MB/s (driver peaks 1574 MB/s) |
| Internal SSD, random 16 KB read, QD1 | 2000 aligned random reads, F_NOCACHE
|
6618 IOPS = 108 MB/s |
| Internal SSD, sequential write | 10.7 GB, 1 MiB blocks, F_NOCACHE
|
641 MB/s |
| Memory immediately reclaimable, idle 8 GB node |
vm_stat free + inactive + speculative |
~2.5 GB (macOS reports 63% "free") |
| External NVMe per node | diskutil list external |
none attached — the 3 GB/s assumption has no hardware behind it yet |
The two numbers that change decisions: the transport is base-RTT-bound, not bandwidth-bound (4 KB costs 645 µs against a 515 µs empty round trip, so the payload is nearly free and UDP buys nothing), and random 16 KB reads at queue depth 1 are 9–11x slower than sequential reads (108 vs 1161 MB/s), which is the entire case for a deep read queue and expert caching. The derived per-token I/O and synchronisation costs are worked out on the Target models page.
Run before a cluster run, on each node. These are read-only checks; expected values are recorded in the Project Tracker as evidence, not in this wiki.
sw_vers # macOS version — identical across nodes
swift --version # toolchain — identical across nodes
xcodebuild -version # Xcode / Metal toolchain
python3 --version # 3.14 across the farm
sysctl -n hw.memsize # unified memory
sysctl -n machdep.cpu.brand_string
df -h / # SSD headroom for this node's shard
ping -c 5 <peer-name> # the name takes the VPN path — measured 1.4–1.8 ms
ping -c 5 <peer-address> # the Ethernet segment — measured 0.49–0.64 msReachability is checked twice on purpose. On this farm the node names resolve over the VPN while the direct segment is reached by address, so a check that only follows names measures the wrong path — and gets a number 2.5x too pessimistic for the all-reduce (R4).
Also verify, once the code exists: identical build version hash, identical model install digest, and that the node can read its own shard at the expected rate.
A second, independent implementation is needed to produce golden traces — bit-exactness has no meaning without a reference (R7, DC-021). Options, none chosen yet:
- a CPU reference path,
- a separate reference implementation run on the developer machine,
- a virtualised or remote reference runner for an independent result.
Virtualisation is available on the farm and on the developer machine (a hypervisor on one Mac, and QEMU elsewhere) if an independent platform is wanted for a cross-check. A virtualised reference must never be the only evidence for a gate: the gate is passed on real hardware.
Source · Issues · Discussions · Sister project: TinyTitan
Project
Design
Development
Sister project: TinyTitan