Skip to content

History / Project Tracker

Revisions

  • Wiki review: remove the dependency framing; flag the plan pages as stale Removed every non-historical 'sister project' reference (Home, Glossary, two Architecture rows, the Roadmap objective, Testbed's dependency section) and the two tracker rows that presupposed adopting another project's code - DC-085 and DC-123. ttd is built from scratch and is MIT, so there is no upstream to track and no code to adopt. Flagged rather than rewritten: Roadmap.md, the milestone pages and most tracker rows still describe the retired DatacenterEngine plan, and Architecture.md's sharding section says expert parallelism while what shipped divides layers. That is a planning decision, not a cleanup.

    @Pummelchen Pummelchen committed Sep 19, 2026
  • DC-138: refresh the next step to D283 and the measured result The row carried the state as of D271. It now records the measured result - a requesting node 6.4-6.6 tok/s and the serving node 1.36 -> 1.64 after two fixes, against a single node's 7.377-7.974 - and moves the next step to the thing actually blocking: --shard-serve-only parses and runs but suppresses the entire run, so the one measurement left about the serving cost (contention or the command-buffer round trip) cannot be taken. Kept lean per docs/task-table-standard.md: what was tried and rejected is in the decision records, and the row points at them.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • wiki: rewrite the project tracker to the task-table standard

    @Pummelchen Pummelchen committed Sep 18, 2026
  • DC-138: the distribution is built, runs, and measures 0.85x a single node The tracker carried projections. The measurement now exists: three repeats, loads recorded, node3 at 1.20-1.47, 48 tokens, 40 slots. A requesting node reached 6.4-6.6 tok/s (medians 6.581 and 6.392) and the serving node 1.36 (median 1.363), against this node's own single-node 7.377-7.974. That is 0.85x and 0.18x, with zero NaN in every run that completed. Three nodes, not four: node4 is the measuring machine and cannot open 192.168.x.x connections. So the ceiling the objective asked to be tested rather than defended was tested, and it is an over-estimate of what this design achieves. Why, measured: the expert read is worth zero (D239), the routed MoE is 19.4 ms of a 137.0 ms step (D221) and is all that can divide, and the serving path costs eight encodes and a blocking wait per request because topK == maxStreamedExperts is a precondition. The single node is the win: 7.377-7.974 against the reference's 7.075 at every generation length. And the three defects along the way were all one class - types and widths asserted by the Metal declaration and unchecked by Swift - each compiling and passing every gate. What would change the answer is not the down-only kernel, because 19.4 of 137.0 cannot become 2.8x however cheaply divided: it is dividing the dense kernels, which no plan here does.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • DC-138: the requesting half is wired and behaviourally verified encodeDecodeRoutedMoE reads outIndices, widens routedX, calls the provider, copies the [D * 8] reply into a persistent buffer and passes it as remotePartials; the CLI supplies the provider from --shard-plan, --shard-node and --shard-peers; both seams are nil by default, so no run without a plan differs from the ones every measurement here was taken on. Verified through the deployed binary rather than reasoned about: a plan with unreachable peers exits 1 instead of falling back silently, a plan without --shard-peers runs at 7.151 tok/s, and the suite is green with productionRoutedPipelineAndHitSplitMatchReference passing. What is not built is the serving half's real Compute, whose trap is recorded before the code: the existing down projection applies the routing weight, so a peer built on it would send w * value and the requester would multiply by w again - a wrong number the bit-exactness contract cannot catch.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • DC-138: the slot count is the lever and the candidate set is not Sweeping slots with the candidate set pinned at 64 by the ownership filter gives a curve lying on top of the un-filtered one at every point: 8 slots 228.4 vs 232.2 ms, 16 180.6 vs 181.2, 24 154.2 vs 154.4, 40 128.3 vs 129.5. Quartering the candidate set changes the step by under 2% while the slot count changes it by 1.78x. Two independent sweeps now agree - D239 varied the filter at fixed slots and found nothing, D243 held the filter while slots varied and found the slot count - and since miss count is a function of both, it is eliminated as well. What remains is the reservation itself: the slots a streamer allocates, fences and recycles. D114 is the precedent for that class of cost, having taken 640 synchronous dispatches to 80 for 1.27x by removing waits rather than work.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • DC-138: the ownership filter tests the read, not the device setOwnedExpertFilter changes which experts are planned into slots, while the phase-1 kernels still run over all eight routed slots with zeros in the unowned ones. Dispatch count and kernel times are unchanged and only the bytes disappear. So D239 establishes that the expert read is not on the critical path, and does not establish that the device work fails to divide - a real sharded node does not run its peers' kernels at all, so its MoE kernel time still falls with the experts it owns. The read term is what D239 removed from the arithmetic; the device term still divides, which leaves the ceiling at about 1.12x for expert sharding and 1.21x with attention sharded as well.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • DC-138: closed by direct measurement, and not in favour of the sweep The tracker had recorded the exposed-IO contradiction as resolved in favour of the cache sweep on D237. D239 overturns that by asking the question directly: with the ownership filter reachable from the CLI, a node reading 64 of 256 experts was measured against one reading all 256 - 128.2 -> 128.3 ms, 133.7 -> 134.5, 128.4 -> 128.5, three alternating pairs, 1.00x / 0.99x / 1.00x. Reading a quarter of the experts changes the step by nothing, so the expert read is not on the critical path and the reachable speedup is whatever the GPU divides: 27.0 ms of a 137.0 ms step, 20%, about 1.14x. D234 and D235 were right; D236 was wrong to withdraw them. And the cache sweep measures neither read volume nor miss count, both having been cut about fourfold by the filter with no effect. D220's 0.521 ms is therefore per cache slot removed rather than per miss, and the causal story under D218, D219 and D220 is wrong even though their measurements stand.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • DC-138: the exposed-IO contradiction is resolved - the counter's nil guard, not a hidden read CommandCompletionClock.latest returns nil unless completionCount == expected, and every call site is 'if let', so a nil skips the increment rather than adding zero. exposed_io_us = 0 therefore means the clock was not ready, not that the read was hidden. That decides the contradiction D236 recorded in favour of the cache sweep: the read is on the critical path, and D234/D235 are withdrawn. D226 and D227's layer-body reading stands, and D228's calibrated projection is back in play at 14.33 tok/s at four nodes without attention sharding and 16.35 with. The rule gains its second half: a counter must be read with its guard, and the value it emits when it cannot count must differ from a legitimate reading.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • DC-138: withdraw the 1.14x conclusion - two of the session's own measurements contradict each other D234 read the engine's exposed_io_us counter as showing the expert read is hidden, and D235 built a ~1.14x ceiling on it. The cache sweep falsifies both: the step falls monotonically from 175 ms at 8 slots to 129 ms at 40, so 156 MiB fewer read per token buys 46 ms and the read is plainly on the critical path. The tracker had already been rewritten with the withdrawn conclusion, so it is restored here to the honest position: the reachable speedup is not settled, the contradiction is the open item, and the next step is to find out what totalExposedIoNanos actually measures. Everything measured and closed is unchanged - the bit-identical exchange at 3.2 ms, the 7.377-7.974 single node above the reference at every length, the 9.02 floor, the 40-slot cache ceiling, the invariant prefetch, and the non-stationary step that makes generation length the fifth required field.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • DC-138: 21 tok/s is not reachable by distributing this engine - the reachable speedup is ~1.14x D235 closes the question the goal opened. Expert sharding divides the expert read; exposed_io_us is zero on all 119 layers traced, so the read is already overlapped and contributes nothing to the critical path at any node count (D234). Dividing a hidden cost recovers nothing. Only the GPU divides, and 56.5% of it is replicated - attention, shared expert, router - against the routed MoE the plan exists to divide. At a measured 137.0 ms step with 62.0 ms of GPU, the divisible part is 27.0 ms, 20% of the step: 1 node 7.30 tok/s, 2 nodes 7.89, 4 nodes 8.34. Attention sharding moves 22.5 ms across and gives 8.80 tok/s, 1.21x. The ceiling the goal asked to test was right, but for a different reason than it was defended with: D183 and D188 reached 1.0-1.1x from an exchange cost 5.4x too high, subtracting the cost of moving data that was never on the critical path. Six independent measurements agree and three initially said otherwise - D217, D219 and D228 projected 15-20 tok/s before the read turned out to be hidden. Also records the closed findings: the 3.2 ms exchange, the 7.377-7.974 single node above the reference at every length, the 9.02 floor, the 40-slot cache ceiling, the invariant prefetch, and the non-stationary step that makes generation length the fifth required field.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • DC-138: the read is already overlapped, so the four-node projection is 14.33 / 16.35 tok/s D226 and D227 concluded the expert read was fully exposed and D228's projection leaned on recovering it. Printing the exposed fraction the engine already computes (D233) shows exposed_io_us is zero on all 119 layers traced, with the overlap clock non-nil exactly on the layers with misses - so the 43.1 ms read is hidden. io_us + wait_us summing to body_us was read as 'the IO happens, then the wait happens'. They do sum to the body; that does not make them sequential. A sum that fits is not a sequence. What the layer spends its time on is the wait: 53.3 ms per token against 61.7-62.8 ms of GPU kernel time, so the wait is the device executing - and a device term divides. The honest four-node projection is 14.33 tok/s without attention sharding and 16.35 with. Also records that the step is not stationary: the body grows 8.5% and the attention kernel 47% from position 16 to 160, making the generation length the fifth field required beside every tok/s claim.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • DC-138: replace the single tok/s figure with the generation-length curve The tracker quoted 7.357-7.451 against the reference's 7.075 as one number. D230 and D232 show that is under-specified: the layer body grows 8.5% and the attention kernel 47% between position 16 and 160, and throughput rises from 32 to 128 tokens as a run's fixed cost is amortised before falling to 256 as the KV grows. Measured across three lengths: 32 tokens 7.377-7.503, 128 tokens 7.815-7.974, 256 tokens 7.546-7.756. All six runs are above the reference's 7.075, so the win no longer depends on the reference having been measured at any particular length - which D231 established is unrecorded. The generation length is now the fifth field required beside every tok/s claim, after the node, its load, the configuration and the repeat count.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • DC-138: 21 tok/s is not reachable by sharding experts; the four-node case is bounded at 10.5-15.4 tok/s Replaces the pre-measurement framing with the measured result, so the next session reads the bound rather than the hypothesis. Measured: the exchange is 3.2 ms per step, not the 17.3 that D183 and D188 inherited (D208); the single node is maxed at 40 slots with a floor of 9.02 tok/s (D218, D219); the step is set by demand misses at 0.521 ms each, 43% of which is already overlapped (D220); and the binding term is replicated GPU at 57% of the kernel time - attention 338.1 ms, shared expert 119.9, router 106.8 - against the routed MoE's 290.5 that the plan exists to divide (D221). Four nodes therefore project 10.51 / 13.31 / 15.35 tok/s by how much of the non-GPU term divides, and even dividing everything else to zero leaves 40.9 ms = 24.4 tok/s against the 47.6 ms that 21 needs. The ~1.0-1.1x ceiling was right by accident: reached from an exchange cost 5.4x too high, surviving because of replicated kernels the earlier decisions never named. The one untried route is to shard attention and the shared expert as well as the experts - attn_norm_qkv is the largest single kernel in the engine and the plan does not divide it, while D93's vocabulary-parallel head is the precedent.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • DC-138: every setting is closed, the exchange is measured at 3.2 ms, and the obstacle is 41.3 ms of host CPU Records the corrected arithmetic so the tracker stops carrying the numbers D208 and D209 replaced. The measured exchange is 3.2 ms per step - not the 17.3 ms that D183, D188, D189, D198, D202 and D203 all subtracted. The single-node step decomposes exactly (132.7 = 49.4 reads + 42.0 device + 41.3 host), and the host share of 31.1% matches the 31-37% of one core measured independently, so the host term is CPU work and does not divide across nodes. With the measured exchange the projection is 72.7 ms -> 13.8 tok/s, an upper bound because the model hides reads the real prefetch ring cannot. Reaching 47.6 ms needs the host loop cut from 41.3 to 16.2 ms. Every setting is exhausted - including threading, which has no knob at all - so what remains is code: spread the host loop across cores (1.74x when this repository did it for its other engine, D94) or cut its dispatch count (D114's shape, 1.27x). Both help the single-node number by the same factor, because the host loop is identical on one node and on four.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • Correct DC-137: the proposed change is already made, and the analogy fails at the stage D201 supersedes this row's original proposal. Issuing layer L's reads before L's attention would repeat a change the repository has already made and measured: RealForwardRunner+Decode.swift:1292 queues the shared MLP before the router readback, because encoding it after left 7.88 ms/token of GPU idle in the attn_tail_router -> shared_expert transition - its own note calls it the largest single component of decode's idle time. That idle is taken and is not part of D198's 97.4 ms. And the analogy does not carry to reads anyway, for a reason worth writing down: the reads cannot be issued before the router readback, because until it returns the CPU does not know which experts the layer wants. That gap is exactly what the prefetch ring exists to bridge, and D200 shows it bridges one read per layer against 1.6 misses. So the row now states the lever that needs no new mechanism: the demand misses are known and are read at 985 MB/s where the same reads achieve 1,940 at depth 1 and 2,940 at depth 8 (D196). The target stays a step time - under 68 ms on node3 - rather than an adjective.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • DC-137: locate the one change the measurements support, with the line and the target D200 identifies the cause and this row makes it actionable. RealForwardRunner+Decode.swift:1384 asks the prefetch ring for what is ready at the moment the MoE needs it, and everything not ready becomes a synchronous demand read. One expert read is 0.91 ms (1,769,472 B at the measured 1,940 MB/s depth-1 rate), one layer's device window is 1.05 ms, and a layer misses 1.6 experts on average - so the ring can hide about one and the rest wait, which is the 97.4 ms of the 132.7 ms step. Inside a layer, attention precedes the MoE and the router that determines the expert set runs in that attention stage, so layer L's reads can be issued before L's attention and collected before L's MoE. D87 describes exactly this shape for the repository's other engine. The target is stated as a step time rather than an adjective: under 68 ms on node3 with the load, configuration and repeat count named, against the present 132.7 ms. D196 bounds the read term at 32.6 ms when issued concurrently, so the composition allows up to 23.8 tok/s - above the 21 the objective asks for.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • DC-136: twelve rounds of read-path measurement, what is closed, and the one factor of two left Records the map so the next session starts from it rather than repeating D187-D197. Measured on node3 with the reference's install at 40 slots: the decode step is 132.7-137.8 ms (7.256-7.537 tok/s over five runs); the expert path reads 91.4-101.0 MiB/token; the disk does ~985 MB/s steady-state, of which the expert path is 73% and ~35 MB/step is unattributed; the device is ~30% occupied and the host 31-37% of one core. The identical demand reads with no compute interleaved run at 1,940 MB/s, and 2,940 at depth 8, measured directly on the install's own layer file with F_NOCACHE. Closed by measurement: speculative decoding and every MTP knob (no draft head in the install); prefetch depth beyond 1, which is 28% WORSE at depth 8 because depth here is speculative and wrong guesses cost more than the concurrency gains; access order, identical scattered or ascending to 0.5%; read size, already one whole expert per pread; and the reader itself, where serial pread is right at this batch size by the repository's own numbers. Open: why the demand reads achieve 985 MB/s where the same reads achieve 1,940 - a question about what the step does between reads. That factor of two is worth about 1.9x on one node if recovered, which is more than every sharding case measured in this session combined, and it is the only route to 21 tok/s the measurements support.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • Open Q11: what is the ~280 MB/s of per-step disk I/O the expert path does not account for D191 and D193 measured the disk at 1,050-1,171 MB/s during a decode whose step is 137.8 ms - about 152 MB per step - against an idle baseline of 0.12-12.9 MB/s on the same node minutes apart. The engine's own TINYTITAN_DECODE_IO_TRACE accounts for 101.0 MiB/token = 106 MB, which is the 64 expert misses a step at 1,769,472 B each. So roughly 46 MB per step, about 30%, is unattributed. Recorded as an open question rather than a conclusion, because this session has already twice drawn a confident inference from a proxy measurement and had to withdraw it: D188 inferred the read was exposed from a byte count divided by a peak rate, and D190 inferred it was free from prefetch being worth only 5%. The candidates here are not distinguishable from the numbers in hand - the prefill's reads landing in the iostat window, the dense payload being re-read rather than held, KV traffic, or windowing over a bursty pattern - and the answer changes which lever is next. It matters in both directions: if the gap is residency it is a removable term, and if it is prefill polluting the window then D193's 820 MB/s achieved rate is itself an underestimate and the read path is nearer its measured 3.44 GB/s ceiling than it appears.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • Close DC-135: the fork is pushed as the distribution branch Pushed to TinyTitan_Datacenter as refs/heads/distribution (a701de7), leaving this repository's MIT main untouched at 4571baa. It is a branch and not a merge because the fork is derived from Apache-2.0 TinyTitan while this repository is MIT - Apache-2.0 code may live in an MIT project only if it travels with the Apache-2.0 text and the upstream NOTICE, and the branch carries both plus section 4(b) change notices on the six pre-existing files it modifies. Two things were found on the way, and the second is the uncomfortable one. The fork was a SHALLOW clone, which made the first push fail with 'did not receive expected object'; and that also means the bundle written earlier and described as recording a complete history was NOT one - git bundle verify said so of the refs it held, and the missing ancestry was simply absent from all of them. git fetch --unshallow fixed both: the push succeeded, and the bundle is now 14.8 MB against the earlier 5.8 MB, which is the difference made visible. A merge into main remains an operator decision and would want third_party/TinyTitan/ extended to cover the derived sources (D100, D124).

    @Pummelchen Pummelchen committed Sep 18, 2026
  • Close DC-134 on its measurement, the target renegotiated by the operator (D186) The operator's decision: accept the measurement and renegotiate the target. So the row closes on D183 rather than staying open against 3x. Measured: the decode step is 139.6 ms and only the mixture divides - 440.3 ms, 13.2% - while 568.4 ms, 17.0%, is replicated work every node repeats. A perfect free four-way division gives 1.11x and the measured 17.3 ms exchange takes it to 0.98x. Stated in the row because the distinction matters: the phase budget is a MEASUREMENT, the four-node ratio is DERIVED from it by arithmetic and NOT measured on four nodes. No four-node run was completed - the call site in encodeDecodeRoutedMoE is unwritten, and the farm cannot hold the 19 GB install on all four nodes, node2 having 12 Gi against node1's 29, node3's 22 and node4's 11. DC-135 stays open and is independent: the 21 fork commits carrying the distribution still have no remote they may legitimately be pushed to.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • DC-134: the farm cannot hold the install on all four nodes The install is 19 GB. node2 has 12 Gi free, against node1's 29, node3's 22 and node4's 11 - and node3 and node4 already hold it, node4 being at 11 Gi because it does. So a four-node run needs about 7 GB freed on node2 first, or an install that is not replicated everywhere, which would change what the run measures. Worth recording next to the renegotiated target because it is a second, independent reason a four-node measurement is not a matter of running the code that now exists.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • Record DC-135: the distributed implementation has no remote it may be pushed to The fork of TinyTitan carrying the distribution work has 21 commits and one remote, Pummelchen/TinyTitan, which AGENTS.md requires be treated as read-only. So the implementation - the shard plan, the exchange frames, the peer channel, the LAN transport, the exact zero-padded fp32 reduce, the participant, the shadow kernel and the integration test that now crosses a real socket - exists on one disk and nowhere else. Mitigation in place, and it is a mitigation rather than a fix: `git bundle create --all` records a COMPLETE history (verified), and the 5.8 MB bundle is on node4 at ~/tt-backup/ and copied to node3's Downloads. Two disks, still one machine apart from the farm's own copies, and a bundle is not a repository anyone can continue from without unpacking it. What would actually settle it is one of two decisions, and both are the operator's: - a remote the fork may legitimately be pushed to, which is the cheap option; or - landing the distribution into this repository, which is the option with an obligation attached - the code originated in an Apache-2.0 project and this repository is MIT, so it travels with attribution (D100, and third_party/TinyTitan/ is the scaffolding that requires it).

    @Pummelchen Pummelchen committed Sep 18, 2026
  • Renegotiate DC-134's target with the measurement that forces it, and mark R2 realised This repository's own rule: a gate that cannot be met is renegotiated in the tracker, with the measurement that forced it. D182 and D183 are that measurement, taken with TINYTITAN_KERNEL_STATS on node3. The decode step is 139.6 ms with the device 33% occupied — but occupancy is the wrong quantity, and separating the roles shows why. Only the mixture divides: moe_phase1_hit 165.0 ms moe_phase1_miss_fixup_phase2 161.8 ms moe_phase1_2_routed 113.5 ms -------- 440.3 ms = 13.2% of the 3350 ms decode wall while these are replicated, identical on every node, and therefore cannot be divided: head_logits 221.1 ms shared_expert 183.7 ms attn_tail_router 163.6 ms -------- 568.4 ms = 17.0% of the wall A perfect, free four-way division of everything that divides gives 139.6 x (0.132/4 + 0.868) = 125.8 ms, which is 7.95 tok/s, or 1.11x. Adding D173's measured 17.3 ms exchange makes it 143.1 ms, or 6.99 tok/s, or 0.98x. To reach 3x, more than 88% of the step would have to divide. It is 13%. So DC-134's target is not reachable by this design on this engine, and the row now says so with the arithmetic rather than leaving it as an unmet ambition. What remains worth doing is named in the row: the vocabulary-parallel head, which is 221.1 ms or 6.6% of the step and is worth MORE than the entire expert exchange — D93 took this repository's own engine from 1.13x to 1.36x with exactly that change, and it is not a sharding change at all but a change to how the head is computed on one node. R2 moves from a risk to realised: the exchange may cost more than the reads it hides, and it does. It is exact and it is on the forward path, but at 13% of the step there is not enough expert work for it to divide, and its 17.3 ms costs more than the division saves.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • Correct the tracker to the measured state: DC-117, DC-128 to DC-133 closed; DC-134 open The tracker was materially out of date. DC-128, DC-129 and DC-130 were recorded Blocked, DC-131 and DC-132 Open, DC-133 Planned, and DC-117 In progress. The measurements that settle them are in docs/repository-decisions D171-D181. DC-117 and DC-128: the install was built, verified (46 files, 20,078,200,501 bytes, every SHA-256 recorded) and measured at 7.30 tok/s (node3 at load 3.6-3.9) and 7.292 (node3 at load 0.95-1.33) against the reference's 7.075, with the engine (Metal TinyTitanCLI) and the slot count (40 = 2.83 GB) named and a median over repeats. The slot sweep reproduces the reference's non-monotonic curve, peaking at 2.83 GB and collapsing at 4.53 GB. DC-129: all four nodes answer. Kept in the row because two findings still matter - node names resolve to Tailscale 100.x while the LAN is 192.168.18.x, so ssh nodeN silently takes the VPN; and node4 cannot open any 192.168.x.x connection under macOS Local Network privacy. DC-130: measured the switch - 118 MB/s and a 0.565 ms median RPC against the Wi-Fi pair's 40 MB/s and 14.992 ms the earlier reasoning rested on. It REVERSED the design: per-step exchange beats per-layer by 2.3x. DC-131: done and verified ON the forward path rather than beside it - ownedExpertFilter at the top of makeExpertCachePlan, reached by encodeDecodeRoutedMoE -> planRoutedExperts -> planExpertsCached, set for every layer by ModelExpertIO.setOwnedExpertFilter. nil is an ABSENT filter, so the single-node path is unchanged. DC-132: the arithmetic is done and proven but the row stays In progress, because the call site is not. ShardReduce is exact by construction - k fixed at 8, fp32 partials, adding zero is exact - so any partition gives the identical bit pattern, and the tests assert bitPattern rather than a tolerance. The kernel takes the peer partials (D180) and remote == nullptr is verified bit-identical (D181). What is missing is the blocking round trip in encodeDecodeRoutedMoE: read moeActs back, exchange, build the [d][8] buffer, pass it. DC-133: done, and it CORRECTED ITS OWN ROW. The 116 MB/node in this row was wrong by 40x, because an expert ID is replicated in every layer: R experts costs R x 40 x 1,769,472 B, so R = 64 is 4.53 GB against a measured 2.83 GB cache inside 8 GB. D179 records it and a test asserts the byte figure so it cannot return quietly. DC-134 stays Open and the goal is left open on it. D179 says the >=3x target must be measured rather than asserted, and D177 says every node on this farm is busy as a standing condition, so a cluster ratio needs a quiet window or must be reported as a bound.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • Tracker rebuilt: open tasks only, in tables, with their DC numbers The operator asked for the tracker to carry open tasks and issues as tables with proper task numbers, and for everything closed or merely informational to be removed. This does that. REMOVED. Every row whose status was Done or Dropped, and the "Closed in this phase" paragraphs that summarised them (DC-107, DC-112, DC-113, DC-115, DC-116, DC-118, DC-119, DC-120, DC-121, DC-122, DC-124, DC-126), the "Later / parked" list - which by its own words held things that are "not tasks until they get a DC-nnn id" - and the thousand-word historical narratives inside individual rows. That history is on News and in the decision records, which is where a reader goes for it; a tracker row is for what is open and what would close it. KEPT, and kept short. 19 existing open rows (DC-008, 051, 052, 053, 060-064, 070-073, 085, 114, 117, 123, 125, 127), each reduced to a task, its dependency, its evidence, and its status. ADDED. Seven rows for the work this repository is actually doing now, which had no DC id at all: the objective moved from grinding this engine to 7 tok/s to taking the sister project's runtime and putting this repository's distribution on top of it (D122, D123, D135). DC-128 build the 4-bit qwen36 install and measure single-node at the reference's own configuration DC-129 farm access - this session is on node4, and node1-node3 refuse authentication DC-130 measure the wired switch, not the Wi-Fi pair the last four rounds were reasoned from DC-131 wire the shard plan into the expert path, filtering the routed set by ownership before planning DC-132 carry the expert exchange over the LAN peer channel DC-133 choose the replicated expert set DC-134 four-node run: bit-identical and measured DC-128 to DC-130 are Blocked and DC-129 is the reason: everything the distribution needs is built and tested, and what stands between it and a four-node run is a login on the other three machines. That is now a row with an ID rather than a paragraph in a chat. Also updated: the Current status block (this node is node4; the farm link is wired - the 0.49-0.64 ms figure is the switch, while every figure in D165-D168 came from macbook-ab over Wi-Fi), the risks table to live risks only, and the open questions to the four that are genuinely open - Q6 is now live rather than theoretical because D166 made expert duplication a requirement rather than a possibility, and Q10 is new. Verified: the table gate passes over the wiki (36 files, 0 misaligned rows), the page carries 26 task rows, and no ID appears twice.

    TinyTitan Datacenter committed Sep 18, 2026
  • Correct the documented Python test count to 428 (D132) D124 changed the provenance checker without its tests; three tests failed from that commit on and were not seen because only the four documentation gates were being run. Fixed, and the suite is 428 now. Also records the int4 head/embedding policy change being reverted as not a standalone change - the fixture install needs regenerating, which belongs with the single rebuild.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • Re-test the decoded-layer cache at the new operating point (D118) D89 rejected it at 5.42 s/step; at 0.626 s/step and load at 68 ms the trade was worth re-asking. 1 GB takes load 68 -> 20 ms -- 70% of the phase -- and the step is still 2% worse (0.626 s against 0.639, three alternated pairs), because mix.gather absorbs 16 ms of it. Default stays 0: the fourth independent confirmation of the memory wall.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • Load four bf16 per lane in the head tile (D117) The head filled a 32x32 tile with one warp-wide 64-byte load per row: 32 instructions for 2 KB, and head measured 12 GB/s on hardware whose memory does ~100. Four bf16 per lane with ushort4 -- eight loads instead of thirty-two -- took head 82 -> 48 ms and the step to 0.628 s, 1.591 tok/s, bit-identical by construction because the tile's contents are unchanged. A load can be perfectly coalesced at the warp level and still instruction-bound; the way to see it is bytes-per-second against the hardware. DC-117 carries the progress.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • Map the dense int4 payloads on the device once (D116) D115 fixed the per-call device copy for the head; the same defect sat in the dense path, called 130 times a step -- 865 MB of copying a step for weights that never change. One bytesNoCopy mapping now covers all three sections and each dispatch addresses them by offset: attn.core 146 -> 118 ms, step 0.678 -> 0.65-0.67 s (about 1.5 tok/s). An expert's row range must never be mapped (those bytes come from the evictable slab cache, so a mapping would dangle), and the key is content-addressed (name#sha256-prefix) because a name is unique within one install and says nothing across two -- D115's bug one level out. DC-117 carries the progress.

    @Pummelchen Pummelchen committed Sep 18, 2026