Repository navigation
Risks and Traps
André Borchert edited this page Sep 18, 2026
·
1 revision
These are the risks and traps and the handling that stands, not open work; open work lives in the Project Tracker.
| ID | Risk | Handling |
|---|---|---|
| R2 | The exchange may cost more than the reads it hides. |
Realised, not merely possible (D182, D183): the exchange is exact and on the forward path, but only 13% of the step is expert work it can divide, so the four-node ceiling is ~1.0-1.1x against a target of 3x, and D173's 17.3 ms costs more than the division saves. DC-051 measured the budget before any scaling claim; DC-130 replaced the Wi-Fi latency the design rested on with the switch's 0.565 ms, and D179 removed a replication cost that was 40x too low |
| R3 | 8 GB per node bounds what is resident. Dense backbone, shared expert, KV state and expert cache all have to fit beside macOS. | Shard policy and model choice are constrained by measured resident cost, not intent. Replication (DC-133) spends this budget directly and must be measured against the 3 GB cache optimum |
| R4 | Transport reality versus documentation. The farm has two paths — 1 Gbit Ethernet at 0.49–0.64 ms and a mesh VPN at 1.4–1.8 ms — and node names resolve over the VPN, so a run that binds what a host name resolves to silently takes the slow path. No Thunderbolt bridge is configured on any node despite two ports each. |
DC-008 records the measurements and the chosen transport. D170 adds: node names resolve to Tailscale addresses (100.x), which is how every command in this session has reached the farm — the LAN path has never been exercised between these nodes |
| R7 |
A guard's configuration is live only when its process is restarted. D140 fixed the disk watchdog's pattern list and the fix was not in the running process; the same class bit twice more (D132, D148). |
The restart is part of the fix, not an afterthought, and a marker is read before it is cleared |
| R9 | Public-repository hygiene. This wiki and the repository are public: credentials, addresses, access paths and model-access keys must never be written down. | Node names are labels (node1…node4) and nothing else identifies a machine; the Testbed page records hardware class, measurements and roles only (DC-083) |
| R10 | Bit-identity across shard counts does not follow from "accumulate in fp32". Pre-summing a node's own experts changes the association order: 11185 of 20000 random top-8 draws (56%) sum differently by 1 ULP. | The reduction must be canonical by construction. Measured in the reference (D154): its reduce is a fixed k = 8 kernel that zero-pads unused slots, so a node contributes into its own slots and the sum order is preserved — which is why the ported plan returns slots rather than a filtered list |
| R11 | bf16 router logits flip the top-k set. Over 20,000 random 256-way routers, 922 (4.61%) changed the top-8 index set. One flipped index diverges the output completely while per-tensor MSE still looks healthy. | Resolved as D5: router logits, normalisation, comparison and top-k run in fp32 with an explicit ascending-expert-id tie-break on both sides. The I3 assertion is a separate test from any numeric tolerance and needs a fixture whose logits sit within 1 ULP, or it cannot fail |
Source · Issues · Discussions · Sister project: TinyTitan
Project
Design
Development
Sister project: TinyTitan