Repository navigation
Project Tracker
Pummelchen edited this page Sep 19, 2026
·
259 revisions
Standard: Task table standard.
| ID | Task | Type | Area | Size | Status | Owner | Next step |
|---|---|---|---|---|---|---|---|
| DC-125 | Add a deterministic regression test for the StreamHandle race that fails without the lock |
Improvement | Tests | S | Open | here | Read one checkpoint's rowsStreaming from many threads and assert every value matches the mapped read, so removing the lock turns the suite red rather than the farm flaky (DC-124) |
| Q9 | Explain the 2,218 expert requests in a five-token prefill (221.8 per token against 320 selections) | Investigation | MoE | M | Open | here | Count distinct experts per layer in the mixture to settle whether de-duplication within a layer explains the count (DC-034) |
| Q8 | Decide whether Core ML / Neural Engine prefill is part of a node's engine and where its sidecar is produced and verified | Investigation | Importer | M | Open | here | Decide and document the prefill boundary and where the sidecar is produced and verified; the conversion toolchain is installed but installing it decides nothing (DC-033, DC-061) |
| Q11 | Account for the ~46 MB/step (~30%) of disk I/O the expert path does not attribute | Investigation | Observability | M | Open | here | Attribute the measured 1,050-1,171 MB/s decode I/O against the trace's 101.0 MiB/token (D191, D193), testing prefill reads in the iostat window, dense re-reads, KV/attention traffic and 2-second windowing as candidates |
| Q10 | Confirm the expert-ownership plan survives contact with real routing and quantify cross-node exchange volume | Investigation | Cluster | M | Open | here | Compare the contiguous 4-node plan (64 of 256 experts each) against the reference's routing statistics and measure how often a routed top-8 expert is not local (DC-131, DC-133) |
| Q6 | Decide whether "unlimited nodes" bounds N by experts-per-layer divided by k or permits expert duplication | Investigation | Cluster | M | Open | here | Measure the replicated expert set and its memory cost at four nodes, now that replication is the lever that reaches throughput rather than a theoretical option (D166, DC-011, DC-050, DC-133) |
| DC-008 | Record the transport ADR's per-hop latency and bandwidth for Thunderbolt bridge, LAN and SFP/QSFP | Investigation | Transport | M | Open | here | Measure latency and bandwidth per hop on the real links (1 Gbit Ethernet 0.49-0.64 ms, mesh VPN 1.4-1.8 ms, no Thunderbolt bridge configured) and table the measured figures instead of the paper's (DC-009) |
| DC-053 | Re-check M1's restated claim with its own instrument and measure the at-least-3x cluster ratio in a quiet window | Investigation | Cluster | M | Open | here | Verify every frozen prompt is identical in both the checkpoint and the install forms, then measure the at-least-3x cluster ratio with the farm quiet (DC-114) |
| DC-051 | Reconcile the per-token synchronisation budget and make the all-reduce receive concurrent | Improvement | Cluster | M | Open | here | Break the ledger's exchange_seconds down against the ff phase that contains it and make the receive concurrent instead of sequential in peer order (measured 15.0% then 71.5/69.5/11.2/63.1% across four nodes, a 6x spread) (DC-050) |
| DC-052 | Tune and record per-node cache budget and prefetch depth at four nodes | Improvement | Streaming | M | Open | here | Record cache budget and prefetch depth per node, resolving why the bit-identical decoded-layer cache loses on an 8 GB node (load is 30.5% of a cached step and is not I/O) and defaults to 0 (DC-050) |
| DC-114 | Overlap the exchange wait with the next layer's decode, holding one decoded layer transiently rather than a whole-model cache | Improvement | Cluster | M | Open | here | Hold one layer of decoded weights across the exchange and measure a cluster step at or below 3.658 s/step with bit-identity intact, or measure that the wait is too small to hide anything (DC-117) |
| DC-132 | Carry the expert exchange over the LAN peer channel on the real forward path | Improvement | Cluster | M | Open | here | Add the blocking round trip in encodeDecodeRoutedMoE (RealForwardRunner+Decode.swift:1773): read moeActs back, exchange, build the [d][8] buffer and pass it, then verify bit-identity on the single-node path (D180, D181) |
| DC-127 | Wire the fused int4 matmul into the expert path so a packed resident slab never materialises as fp32, the precondition for residency | Improvement | MoE | L | Open | here | Extend WeightSource to expose the packed form and batch the per-layer fused int4 expert matmuls (640 synchronous dispatches remain), then measure one binary alternated for a lower step time with the trace digest unchanged (D108-D110) |
| DC-136 | Recover the factor of two between the 985 MB/s demand reads and the 1,940 MB/s the same reads reach, to approach 21 tok/s | Improvement | Streaming | L | Open | here | Make the demand reads overlap the step's compute, or measure where the half is lost, toward a step under 68 ms on node3 with load, configuration and repeat count named; the read-path levers (MTP, prefetch depth beyond 1, access order, read size, serial pread) are closed by measurement (D187-D197, D196, D193) |
| DC-137 | Read the demand expert misses faster so the 985 MB/s demand path approaches the measured 1,940 MB/s | Improvement | Streaming | L | Open | here | Read the known demand set faster, since issuing reads before attention is already taken (RealForwardRunner+Decode.swift:1292) and the router readback is the hard ordering, toward a step under 68 ms on node3 with load, configuration and repeat count named, against the present 132.7 ms (D196, D198, D201) |
| DC-138 | Close the gap from the built distribution's measured 0.85x to the 2.8x the target needs (requesting node 6.4-6.6 tok/s, serving 1.36->1.64 after two fixes, against a single node's 7.377-7.974) | Improvement | Cluster | L | Open | here | Fix --shard-serve-only, which parses and runs but suppresses the entire run - with it a node prints nothing at all, not even the shard lines preceding the serve block (D283) - then measure an idle serving node to settle whether its 2.95 ms a request is contention or the command-buffer round trip. The route past ~1.0x is tensor-parallel sharding of the dense kernels, which ShardPlan has no concept of (D239, D221, D272) |
| DC-063 | Check sparse-attention block selection for exactness against the reference | Chore | Attention | M | Parked | here | Revive when DC-061 lands: assert selections match the reference exactly, not approximately |
| DC-060 | Add the DeepSeek importer with transcoding of native FP4 experts and FP8 dense weights | Improvement | Importer | L | Parked | here | Revive when P5/M4 opens (after DC-064 is gate-ready): transcode rather than requantize so no requantization step exists in the path (DC-045) |
| DC-061 | Implement compressed / sparse hybrid attention and constant-size recurrent state | Improvement | Attention | L | Parked | here | Revive when DC-060 lands: run the model at the reference context length |
| DC-062 | Reach MTP speculative decoding parity with the reference protocol | Improvement | Decode | L | Parked | here | Revive when DC-060 lands: make draft and verify match the reference protocol |
| DC-064 | Pass the M4 gate: match the reference at 128K context | Investigation | Cluster | L | Parked | here | Revive when DC-063 lands: run the exactness protocol at 128K and record the gate measurement |
| DC-070 | Add the Qwen3.8-Flash-Next importer and match the reference | Improvement | Importer | L | Parked | here | Revive when DC-064 (M4) passes: run the same exactness protocol as M4 |
| DC-071 | Handle the sparse n-gram lookup table with correct per-node sharing | Improvement | MoE | L | Parked | here | Revive when DC-070 lands: verify table lookups are correct and bounded in memory per node |
| DC-072 | Reproduce or explain M3's expert-parallel scaling on the 125B-A6B shape | Improvement | Cluster | L | Parked | here | Revive when DC-070 lands: reproduce gate M3's ratio on this shape or explain it |
| DC-073 | Pass the M5 gate: match the reference on the same protocol as M4 | Investigation | Cluster | L | Parked | here | Revive when DC-072 lands: record the M5 measurement |
Source · Issues · Discussions · Sister project: TinyTitan
Project
Design
Development
Sister project: TinyTitan