Skip to content

Project Tracker

Pummelchen edited this page Sep 19, 2026 · 259 revisions

Project Tracker

Standard: Task table standard.

Tasks

ID Task Type Area Size Status Owner Next step
DC-125 Add a deterministic regression test for the StreamHandle race that fails without the lock Improvement Tests S Open here Read one checkpoint's rowsStreaming from many threads and assert every value matches the mapped read, so removing the lock turns the suite red rather than the farm flaky (DC-124)
Q9 Explain the 2,218 expert requests in a five-token prefill (221.8 per token against 320 selections) Investigation MoE M Open here Count distinct experts per layer in the mixture to settle whether de-duplication within a layer explains the count (DC-034)
Q8 Decide whether Core ML / Neural Engine prefill is part of a node's engine and where its sidecar is produced and verified Investigation Importer M Open here Decide and document the prefill boundary and where the sidecar is produced and verified; the conversion toolchain is installed but installing it decides nothing (DC-033, DC-061)
Q11 Account for the ~46 MB/step (~30%) of disk I/O the expert path does not attribute Investigation Observability M Open here Attribute the measured 1,050-1,171 MB/s decode I/O against the trace's 101.0 MiB/token (D191, D193), testing prefill reads in the iostat window, dense re-reads, KV/attention traffic and 2-second windowing as candidates
Q10 Confirm the expert-ownership plan survives contact with real routing and quantify cross-node exchange volume Investigation Cluster M Open here Compare the contiguous 4-node plan (64 of 256 experts each) against the reference's routing statistics and measure how often a routed top-8 expert is not local (DC-131, DC-133)
Q6 Decide whether "unlimited nodes" bounds N by experts-per-layer divided by k or permits expert duplication Investigation Cluster M Open here Measure the replicated expert set and its memory cost at four nodes, now that replication is the lever that reaches throughput rather than a theoretical option (D166, DC-011, DC-050, DC-133)
DC-008 Record the transport ADR's per-hop latency and bandwidth for Thunderbolt bridge, LAN and SFP/QSFP Investigation Transport M Open here Measure latency and bandwidth per hop on the real links (1 Gbit Ethernet 0.49-0.64 ms, mesh VPN 1.4-1.8 ms, no Thunderbolt bridge configured) and table the measured figures instead of the paper's (DC-009)
DC-053 Re-check M1's restated claim with its own instrument and measure the at-least-3x cluster ratio in a quiet window Investigation Cluster M Open here Verify every frozen prompt is identical in both the checkpoint and the install forms, then measure the at-least-3x cluster ratio with the farm quiet (DC-114)
DC-051 Reconcile the per-token synchronisation budget and make the all-reduce receive concurrent Improvement Cluster M Open here Break the ledger's exchange_seconds down against the ff phase that contains it and make the receive concurrent instead of sequential in peer order (measured 15.0% then 71.5/69.5/11.2/63.1% across four nodes, a 6x spread) (DC-050)
DC-052 Tune and record per-node cache budget and prefetch depth at four nodes Improvement Streaming M Open here Record cache budget and prefetch depth per node, resolving why the bit-identical decoded-layer cache loses on an 8 GB node (load is 30.5% of a cached step and is not I/O) and defaults to 0 (DC-050)
DC-114 Overlap the exchange wait with the next layer's decode, holding one decoded layer transiently rather than a whole-model cache Improvement Cluster M Open here Hold one layer of decoded weights across the exchange and measure a cluster step at or below 3.658 s/step with bit-identity intact, or measure that the wait is too small to hide anything (DC-117)
DC-132 Carry the expert exchange over the LAN peer channel on the real forward path Improvement Cluster M Open here Add the blocking round trip in encodeDecodeRoutedMoE (RealForwardRunner+Decode.swift:1773): read moeActs back, exchange, build the [d][8] buffer and pass it, then verify bit-identity on the single-node path (D180, D181)
DC-127 Wire the fused int4 matmul into the expert path so a packed resident slab never materialises as fp32, the precondition for residency Improvement MoE L Open here Extend WeightSource to expose the packed form and batch the per-layer fused int4 expert matmuls (640 synchronous dispatches remain), then measure one binary alternated for a lower step time with the trace digest unchanged (D108-D110)
DC-136 Recover the factor of two between the 985 MB/s demand reads and the 1,940 MB/s the same reads reach, to approach 21 tok/s Improvement Streaming L Open here Make the demand reads overlap the step's compute, or measure where the half is lost, toward a step under 68 ms on node3 with load, configuration and repeat count named; the read-path levers (MTP, prefetch depth beyond 1, access order, read size, serial pread) are closed by measurement (D187-D197, D196, D193)
DC-137 Read the demand expert misses faster so the 985 MB/s demand path approaches the measured 1,940 MB/s Improvement Streaming L Open here Read the known demand set faster, since issuing reads before attention is already taken (RealForwardRunner+Decode.swift:1292) and the router readback is the hard ordering, toward a step under 68 ms on node3 with load, configuration and repeat count named, against the present 132.7 ms (D196, D198, D201)
DC-138 Close the gap from the built distribution's measured 0.85x to the 2.8x the target needs (requesting node 6.4-6.6 tok/s, serving 1.36->1.64 after two fixes, against a single node's 7.377-7.974) Improvement Cluster L Open here Fix --shard-serve-only, which parses and runs but suppresses the entire run - with it a node prints nothing at all, not even the shard lines preceding the serve block (D283) - then measure an idle serving node to settle whether its 2.95 ms a request is contention or the command-buffer round trip. The route past ~1.0x is tensor-parallel sharding of the dense kernels, which ShardPlan has no concept of (D239, D221, D272)
DC-063 Check sparse-attention block selection for exactness against the reference Chore Attention M Parked here Revive when DC-061 lands: assert selections match the reference exactly, not approximately
DC-060 Add the DeepSeek importer with transcoding of native FP4 experts and FP8 dense weights Improvement Importer L Parked here Revive when P5/M4 opens (after DC-064 is gate-ready): transcode rather than requantize so no requantization step exists in the path (DC-045)
DC-061 Implement compressed / sparse hybrid attention and constant-size recurrent state Improvement Attention L Parked here Revive when DC-060 lands: run the model at the reference context length
DC-062 Reach MTP speculative decoding parity with the reference protocol Improvement Decode L Parked here Revive when DC-060 lands: make draft and verify match the reference protocol
DC-064 Pass the M4 gate: match the reference at 128K context Investigation Cluster L Parked here Revive when DC-063 lands: run the exactness protocol at 128K and record the gate measurement
DC-070 Add the Qwen3.8-Flash-Next importer and match the reference Improvement Importer L Parked here Revive when DC-064 (M4) passes: run the same exactness protocol as M4
DC-071 Handle the sparse n-gram lookup table with correct per-node sharing Improvement MoE L Parked here Revive when DC-070 lands: verify table lookups are correct and bounded in memory per node
DC-072 Reproduce or explain M3's expert-parallel scaling on the 125B-A6B shape Improvement Cluster L Parked here Revive when DC-070 lands: reproduce gate M3's ratio on this shape or explain it
DC-073 Pass the M5 gate: match the reference on the same protocol as M4 Investigation Cluster L Parked here Revive when DC-072 lands: record the M5 measurement
TinyTitan Datacenter

Project

Design

Development

Sister project: TinyTitan

Clone this wiki locally