News: paused, with the handover rewritten for TinyTitan Datacenter after 1.1
News: all four nodes measured alone and as a chain on the fixed binary
News: the four-node chain runs, and 6.143 tok/s confirms the released 6.016
News: two-stage chain validated, four-stage still failing, no cluster number
News: the single-node target is quantified as unreachable on this hardware 141.6 ms/token = 66.5 I/O + 71.3 GPU + 3.8 encode. GPU plus encode alone is 75.1 ms, or 13.3 tok/s, so ~13 is the ceiling only if expert I/O becomes entirely free - 1.8 ms of room against 66.5, a 97% elimination. It cannot be eliminated that far here: the cache is RAM-capped at 40 slots, concurrency does not help at any level, two processes split one serial SSD, and there is no batch dimension with one forward step process-wide. Closed with the number unmet rather than left open.
News: the four-stage figure's command was never recorded, so it cannot be checked The release's headline cluster number, 6.016 tok/s, cannot be re-derived from the record: searching the wiki and docs for the stage wiring returns nothing, so whether it ran with the cache named is unknowable. That is a reporting defect against this project's own standard, which asks for the exact command so a result can be re-derived. It does not make the figure wrong - the per-node numbers are confirmed twice - but it makes the cluster figure unverifiable, and the next chain run must record its command and cache size.
News: withdraw the published full-overlap claim from the release notes and the live release body
News: the fix is neutral for the released per-node figures; the four-stage figure stays open At the release settings the CLI with no cache flag gives 7.998 and 8.008 tok/s against 8.002 and 7.916 with the explicit 40 it was released with - 8.003 against 7.959, so the default now lands where the release aimed and the per-node figures need no correction. They were always taken with the cache named, which is why the defect hid behind them. The four-stage 6.016 figure is not yet established as taken that way, so it may be understated and is settled by re-running the chain.
News: the page-cache inversion does not reproduce, and the reader win is re-verified The budget's own note claims 4-bit measured 13.61 tok/s at 16 slots against 8.78 at 128 under the page-cache policy. Swept: 8 / 16 / 24 / 32 / 40 slots give 6.319 / 6.769 / 7.395 / 8.138 / 8.552 - monotonic, opposite in sign, so the documented configuration loses 21% and the target is not reachable that way. The note is the stated reason the budget constant exists, so this is a documentation defect as well as a negative. Re-verified under the corrected budget, the serial reader beats the parallel default by 8.334 against 7.717 (+8.0%, no overlap), up from +5.1%.
News: the budget fix reaches the CLI default path too No cache flag now gives 7.740 and 7.779 tok/s over 64 tokens, the same band as every explicit-40 figure here - which is the point, since those were always taken with the flag passed by hand and the default path had never been measured. Widths: ~7.76 CLI default, 9.27-9.34 server, target ~13 unreached, with the proposed mechanism ruled out and memory budget the lever that paid.
News: cache budget fixed to a third of physical memory, 4.1x recovered
News: the expert-cache budget is a fixed 8 GiB and ignores the machine defaultExpertCacheBudgetBytes is 8 << 30, divided by the per-slot footprint and snapped. On a box with ~4.5 GB usable that licenses a cache that cannot fit and the process pages - the same failure the function's own comment documents for 8-bit on a 24 GiB machine, now firing for 4-bit on an 8 GB mini. The fix is to derive the budget from usable physical memory rather than lower the constant.
News: the server's automatic expert-cache selection costs 4x Same server, prompt and method; only the cache size named. Automatic selection: 8 tokens in 9.20s and 48 in 26.86s (2.26 tok/s). With --expert-cache-slots 40: 8 in 5.48s and 48 in 9.76s (9.34 tok/s). The earlier 'about 3.5x slower than the CLI' was really a 4x smaller cache. 9.34 tok/s is the highest single-node figure measured, which puts the target ~1.4x away rather than 2x and makes cache size, not overlap, the lever. Worth checking the CLI's own default, since every CLI figure here passed the flag explicitly.
News: correct the server throughput figures, which included prefill Total request time over completion tokens counts prefill as generation, and the stopping prompt made the error large. The marginal rate at two request sizes is 40 tokens in 17.67 s = 2.26 tok/s. The direction survives - the server is about 3.5x the CLI's decode time - but the numbers and the concurrency scaling computed from them are withdrawn. Width is not the cause; the expert-cache default is the candidate.
News: server template fix, and concurrency does not pay on this hardware
News: concurrent-sequence measurement blocked by a missing chat template The server will not load the installed model - installed tokenizer is missing chat_template.jinja - while the CLI runs it because a raw prompt needs no template. Reported rather than repaired: the project's rule is that an unrunnable check is not checked and never fixed by reinstalling a model. The four-sequence aggregate prediction therefore stays unmeasured.
News: the server already runs at width four; only the gate stops the speed-up maxConcurrentSequences defaults to 4 and sessionSlots - the width the runner and scratch are built with - derives from it, so four sequences can be admitted and held. ForwardStepGate only stops their forward passes interleaving. Because the step leaves the GPU idle 66.5 of 141.6 ms, four admitted sequences time-sliced should already give ~28 tok/s aggregate on one node with no new code; true batching is what would raise the rate per sequence.
News: the needed overlap is gated off, and there is no batch dimension No batchSize, sequenceCount or decodeBatch exists in the runtime and the runner holds no sequence array. The server admits several generations but ForwardStepGate - a busy flag and a FIFO waiter list, acquire() waiting until no step is in flight - keeps their forward passes from interleaving, so one generation's read can never overlap another's compute. That is the remaining route to the target and it is deliberately closed. It needs a batch dimension, not a faster disk.
News: the serial reader's win is not cache warmth, and the device does not reward concurrency Three consecutive runs per reader: serial 7.998/8.042/8.036, parallel 7.832/7.793/7.745 - both flat. A page-cache explanation requires warm-up; there is none. A serial reader beating a four-thread one is the third independent sign that concurrency does not help this device, after one worker matching sixteen and a constant per-miss cost across a 65% change in miss count.
News: the parallel cache-bypassing reader costs 5%, and worker count is not the bound Alternating the two readers three runs each: serial 8.101/8.102/8.133 tok/s against parallel 7.687/7.744/7.729 - +5.1% with no overlap, and await down 100 ms. bypassCache sends every read to the device while the serial path reads through the page cache. Separately, ExpertIOScheduler's worker count is now TINYTITAN_EXPERT_IO_WORKERS, and 1 worker matches 16, so the reads were never concurrent and the serialisation is below the scheduler.
News: hits and misses are already split; lead time is the untested axis The decode path keeps a committed phase-1 hit command buffer, separate hit and miss slot lists and an early miss fetch, so the structure to compute resident experts under the missing one exists and is not paying. Await over per-read cost is a constant miss count (66 vs 40 misses at 16 vs 40 cache slots, ~1.65 ms each), so reads pay full latency individually. prefetchAhead accepts only 1 or 2 and the ring holds ~2 slots, so more than two reads can never be outstanding; depth and slots have never been raised together.
News: the cache route is RAM-capped, so it cannot reach the target Peak RSS by cache slots: 3.33 / 4.83 / 4.97 / 5.14 GB at 16 / 40 / 48 / 64 - about 62 MB per slot, with 40 slots already at 4.83 GB against ~4.5 GB usable. Swap is flat at 40, rises at 48 and explodes at 64, and throughput follows it down. The non-cache floor is ~2.3 GB, leaving ~2.2 GB for experts, which is the 40-slot optimum by arithmetic. Freeing 1.15 GB would buy ~18 slots - about 8 tok/s, not 13. A second token in flight is what remains.
News: knobs exhausted, and EARLY_HITS shows the two phases are one critical path Five profiled runs spread 0.75% - noise. But EARLY_HITS moved time between the buckets: await up 2069 -> 2172 ms, GPU waits down 2128 -> 2001 ms, throughput unchanged. A knob that trades one bucket for the other one-for-one with no net effect is evidence they are a single serial path, which is what the cache sweep said. Overlap needs a second token in flight; the only remaining lever with measured headroom is the expert cache.
News: the ring is worth 15%, and precision hurts Ring on/off pairs on one node: 7.385 and 7.437 with it, 6.357 and 6.428 without - about a tok/s - and it removes ~600 ms of the 2660 ms expert-I/O await. But it adopts only 55% of what it issues, and sweeping the precision margin to 78% precision costs 0.56 tok/s while the await rises. The engine's comment that wrong reads steal demand reads' device time is backwards: coverage is the constraint, not precision.
News: the expert read is serialised, and exposed_io is wrong about it Two profiled one-node runs differing only in cache size: 16 slots gives 184.4 ms/token with 108.6 ms/token of expert-I/O await, 40 slots gives 141.6 with 66.1. The deltas are -42.8 and -42.5 ms - a ratio of 1.007 - so the I/O is on the critical path and removing it removes step time one for one. totalExposedIoNanos reports 0.0 ms in both runs, so it is not answering whether the I/O adds to the step. The earlier 'the overlap already works' conclusion is withdrawn; the buckets summing to the step was the arithmetic that already said so. The prize is ~1.9x on one node.
Wiki review: remove the dependency framing; flag the plan pages as stale Removed every non-historical 'sister project' reference (Home, Glossary, two Architecture rows, the Roadmap objective, Testbed's dependency section) and the two tracker rows that presupposed adopting another project's code - DC-085 and DC-123. ttd is built from scratch and is MIT, so there is no upstream to track and no code to adopt. Flagged rather than rewritten: Roadmap.md, the milestone pages and most tracker rows still describe the retired DatacenterEngine plan, and Architecture.md's sharding section says expert parallelism while what shipped divides layers. That is a planning decision, not a cleanup.
News: ttd is not based on turbo-fieldfare
News: the project is ttd, a dedicated repository
Home: drop the sister-project row, so this wiki no longer points at the runtime repository
News: main now carries the measured engine; the previous main is retired to a tag main is the 28-target streaming runtime that measured 7.146-7.937 tok/s per node (mean 7.73) and 6.016 tok/s through the four-stage chain. The DatacenterEngine line at 55d84fa is preserved as the annotated tag retired-datacenter-engine rather than deleted, the distribution branch is gone, and only main remains.