Skip to content

Risks and Traps

André Borchert edited this page Sep 18, 2026 · 1 revision

Risks and traps

These are the risks and traps and the handling that stands, not open work; open work lives in the Project Tracker.

ID Risk Handling
R2 The exchange may cost more than the reads it hides. Realised, not merely possible (D182, D183): the exchange is exact and on the forward path, but only 13% of the step is expert work it can divide, so the four-node ceiling is ~1.0-1.1x against a target of 3x, and D173's 17.3 ms costs more than the division saves. DC-051 measured the budget before any scaling claim; DC-130 replaced the Wi-Fi latency the design rested on with the switch's 0.565 ms, and D179 removed a replication cost that was 40x too low
R3 8 GB per node bounds what is resident. Dense backbone, shared expert, KV state and expert cache all have to fit beside macOS. Shard policy and model choice are constrained by measured resident cost, not intent. Replication (DC-133) spends this budget directly and must be measured against the 3 GB cache optimum
R4 Transport reality versus documentation. The farm has two paths — 1 Gbit Ethernet at 0.49–0.64 ms and a mesh VPN at 1.4–1.8 ms — and node names resolve over the VPN, so a run that binds what a host name resolves to silently takes the slow path. No Thunderbolt bridge is configured on any node despite two ports each. DC-008 records the measurements and the chosen transport. D170 adds: node names resolve to Tailscale addresses (100.x), which is how every command in this session has reached the farm — the LAN path has never been exercised between these nodes
R7 A guard's configuration is live only when its process is restarted. D140 fixed the disk watchdog's pattern list and the fix was not in the running process; the same class bit twice more (D132, D148). The restart is part of the fix, not an afterthought, and a marker is read before it is cleared
R9 Public-repository hygiene. This wiki and the repository are public: credentials, addresses, access paths and model-access keys must never be written down. Node names are labels (node1…node4) and nothing else identifies a machine; the Testbed page records hardware class, measurements and roles only (DC-083)
R10 Bit-identity across shard counts does not follow from "accumulate in fp32". Pre-summing a node's own experts changes the association order: 11185 of 20000 random top-8 draws (56%) sum differently by 1 ULP. The reduction must be canonical by construction. Measured in the reference (D154): its reduce is a fixed k = 8 kernel that zero-pads unused slots, so a node contributes into its own slots and the sum order is preserved — which is why the ported plan returns slots rather than a filtered list
R11 bf16 router logits flip the top-k set. Over 20,000 random 256-way routers, 922 (4.61%) changed the top-8 index set. One flipped index diverges the output completely while per-tensor MSE still looks healthy. Resolved as D5: router logits, normalisation, comparison and top-k run in fp32 with an explicit ascending-expert-id tie-break on both sides. The I3 assertion is a separate test from any numeric tolerance and needs a fixture whose logits sit within 1 ULP, or it cannot fail
TinyTitan Datacenter

Project

Design

Development

Sister project: TinyTitan

Clone this wiki locally