Skip to content

History / Project Tracker

Revisions

  • Map the head on the device once instead of copying it every step (D115) MetalBf16Matmul.matmul copies its whole weight per call and the head is 1.017 GB in 31 blocks -- 1.04 GB of copying a step for ~1 ms of arithmetic. Mapped once with makeBuffer(bytesNoCopy:) over the payload the engine already holds: head 116 -> 82 ms, +18% on the step (0.757 s against 0.896 in four alternated pairs), default 0.678 s = 1.476 tok/s. The key must name the (tensor, row window) pair -- on the name alone node 1 reused node 0's head and the sharded test produced the wrong tokens. And the read wall is now measured from four sides: the device does 1.65 GB/s cold and the preload already beats it at 1.82 (20 GB/s warm); a large expert cache does find cross-step reuse (43% at 1 GiB, 53% at 1.5) but costs more in memory pressure than it saves in every configuration tried, including with the page cache disabled. 7 tok/s is 143 ms/step against a 238 ms read and 440 ms of everything else. DC-117 carries the progress.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • Encode a layer's experts into one command buffer (D114) Each of the 640 fused calls a decode step made built a command buffer, committed it and blocked on it, and at these shapes the kernel is microseconds of a ~0.3 ms call -- the wait was the cost. One command buffer per layer instead of one wait per expert: mix.read 179 -> 84 ms, mix.down 151 -> 52 ms, step 0.919 -> 0.722 s (1.089 -> 1.384 tok/s). The batch is taken only when every chosen expert was routed the same number of rows (every decode step, not every prefill). The bug worth remembering: the batch's 40-buffer request shared a grow-only buffer cache with the dense projections' 5-buffer request, and that cache reallocates its whole set when the count changes -- 2.47 s/step against 0.92, showing in mix.gather, a phase the change never touched. DC-117 carries the progress.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • Fan the preload out per projection, and record that a partial payload cache is catastrophic (D113) A slab is three preads, so fanning over eight experts gave eight threads six sequential reads each; one task per (expert, projection) is the same bytes with twice the requests in flight, and five runs moved the median 0.984 -> 0.919 s (1.016 -> 1.089 tok/s). The warning is the other measurement: 256 MiB of payload cache gives 2.34 s/step and 1024 gives 2.58, because the LRU holds part of the dense payload and not the head, so both re-read every step -- zero is better than either (1.359 s) because it streams uniformly. 2048 MiB is near-minimal and load-bearing, not a knob. DC-117 carries the progress.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • Multiply the attention projections from their stored form, and shrink the slab cache (D112) The four int4 attention projections are 1.02 GB of fp32 a step and now multiply from their stored form like the GDN three: load 119 -> 54 ms, step 1.087 -> 1.025 s. And with the buffer cache on (D111), the slab cache's optimum moved down -- three alternated pairs put 128 MiB at 0.930 s against 256 at 0.957, with 64/96/128/192 flat between 96 and 128 -- while zero is worse than every non-zero size (1.595 s), because preloadPacked declines with no cache and the read fan-out disappears. Default 128 MiB. 0.920 -> 1.016 tok/s, peak RSS 2.57-2.67 GB, 63% of the step now the expert path. DC-117 carries the progress.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • Read the install through the page cache, and stop decoding the dense int4 projections (D111) The install can now be read through the kernel's buffer cache (SHARD_INSTALL_CACHED, on by default) -- right for a repeating 582 MB working set even though it was wrong for the 20 GB sequential scan D58 records: cold mix.gather 469 ms, warm 266-286, peak RSS/free disk/swap unchanged. A clean cold A/B is not available (purge needs root, and a warm cache serves the F_NOCACHE arm too), so the evidence is the cold/warm contrast and that the switch is what populates it. And the three Gated DeltaNet int4 projections -- 581 MB of the install's 738 MB of dense int4 -- now come from their stored form through the fused kernel, so the fp32 array is never built: load 350 -> 119 ms, digest unchanged. The bug between them: the first packedTensor used the row-range reader and sent 7.5 GB a step back to the device for bytes the payload cache already held, visible only in the byte counter (reads 8.46 -> 14.28 GB). 0.682 -> 0.920 tok/s over the round. DC-117 carries the progress.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • The int4 kernel was loading one byte at a time -- fixing that made the packed expert path pay (D110) The fused expert kernel measured 4.4 GB/s on memory that does ~100: lane l and lane l+1 walked rows 1 KB apart and every element was its own uchar load. Reading each row as a uint4 (32 codes) made it 3.6-4x faster (32768x2048: 8.734 -> 2.196 ms), and the grid caught a trap on the way: #pragma unroll let .relaxed reassociate the accumulator chain and moved one output by 1 ULP, so the unroll is gone and the pipeline is .safe. The packed slab cache was swept and smaller is better -- 256 MiB 1.465 s, 512 1.467, 768 1.483, 1024 1.667 -- because the resident bytes cost more in memory pressure than the reads they save (D106 a third time). Five alternated pairs: fused 1.465 s against the split path's 1.539, digest unchanged, at lower peak RSS. Both defaults are now the measured ones. Three suite-found bugs fixed rather than excused: a 256-byte default budget that made the cache look empty, a preload that chose its destination by switch instead of capability, and a fused counter that reported 13,824 elements read on a forward served entirely from memory. 0.638 -> 0.682 tok/s. DC-117 and DC-127 carry the progress.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • The LM head comes off the device every step -- and three of the round's ideas were losses (D109) A fresh profile put the decode step at 1769 ms (mix.read 805, load 361, head 255, attn.core 202) and 0.565 tok/s. LOST AND RECORDED AS LOST: the GPU unpack for the dense tensor(named:) path (load 353 -> 534 ms); the fused int4 expert path, a wash (2.29 s/step with no slab cache, 1.575 with a 1 GB one against a 1.69 s control -- 640 per-slab dispatch-and-wait pairs eat what the kernel saves); and the first head-residency attempt, which took the node into D106's swap failure twice (31 concurrent blocks each reading the whole 1.017 GB, and UncachedFile.readData building every read twice -- 2.03 GB transient for one 1.017 GB block). Both bugs are fixed. KEPT: the LM head on the GPU from its stored bf16 (MetalBf16Matmul) -- head 255 -> 118 ms, step 1.718 -> 1.580 s in an alternated A/B, digest unchanged -- plus three small measured fixes. Default now 1.567 s/step, 0.638 tok/s, peak RSS 3.10 GB, 266 tests, 0 failures. Tests 259 -> 266; DC-117 and DC-127 carry the D109 progress; open tasks unchanged.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • The fused int4 matmul works on the GPU and is 1.2-2.4x faster -- and the unpack was never the bottleneck (D108) MetalInt4Matmul dequantises inside the matmul, so the ~6.5 GB of Float a token materialises never exists. It is bit-identical to dequantizeInt4 followed by Ops.orderedMatmul over the D107 grid, and in release it beats unpack-then-matmul 1.23x / 1.45x / 2.36x on the real gate_up, down and head-width shapes -- 0.114 s of a 1.74 s step, about 6.6%, the same order as the 7% the CPU probe bounded. mix.read is the disk read, so the route to 7 tok/s is residency, and this kernel is its precondition. The grid also found that an Apple GPU flushes a denormal product to zero where the CPU keeps it, for this kernel and for the MetalMatmul already in the tree, and no math mode changes it -- now a named test. Home and Tests: 253 -> 259 Swift tests. Tracker: DC-113 and DC-117 carry the D108 progress, DC-127 is the wiring row (the provider must hand over a packed payload, and the slot bank stops holding fp32), and open tasks 35 -> 36.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • The fused int4 kernel is 4.7x slower: every CPU-side lever is now measured and closed (D107) 0.33 ms split against 1.55 ms fused in release (0.14x in debug), bit-identical so speed was the only question -- and the probe bounded the prize at ~0.12 s/step (~7%), so there was almost nothing to win. Reverted on the DC-122 precedent. With D105 and D106 this closes the CPU phase: two levers kept, three measured and refused, ~0.55 tok/s. DC-113 is promoted from candidate to route -- packed int4 weights in GPU buffers with the kernel dequantising.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • The packed slab cache works and loses: this node's constraint is memory (D106) 45% hits and 1.08 -> 0.65 GB/step at 1 GB of cache, mix.read 0.80 -> 0.54 s -- and the step gets WORSE (1.931-2.022 against 1.735) because load nearly doubles and head grows on a machine swapping 1.94 of 3.07 GB. The fourth mechanism to reach that verdict after D89, D98 and D105, which makes the pattern the finding: every cache-more route on an 8 GB Mac mini is now measured and closed. Default 0, SHARD_SLAB_CACHE_MB is the knob. The remaining CPU lever is not materialising fp32 at all (DC-120), which costs no residency.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • The predictive prefetch is built and it loses -- because the reads are saturated (D105) On 1.859/1.816 s/step against off 1.854/1.721, with mix.read LARGER when it is on (0.830/0.813 vs 0.790/0.760) and the digest identical. D101's fan-out removed the latency problem, so the device is saturated and a second read stream takes bandwidth from the first. SHARD_EXPERT_PREDICT=1 enables it; the default is off. DC-120's direction is sharpened: the bank holds [Float], so 512 MB buys ~61 of the ~773 slices a token asks for -- the same bytes packed as int4 buy four times as many, which is the next lever.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • The contract matmul is threaded on width-aligned chunks: 0.443 -> ~0.55 tok/s (D104) Every chunk boundary is a multiple of four, so a non-final chunk covers whole four-wide groups with no tail and the final chunk performs the serial body's own out % 4 tail -- the rule the first attempt was missing. attn.core 0.493/0.495 -> 0.209/0.202, mix.gateup 0.224/0.225 -> 0.077/0.073, mix.down 0.102 -> 0.049/0.047, step 4.124/4.103 -> 1.866/1.797 s, digest identical in all four arms. DC-122 is done, both halves.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • DC-124 found and fixed: a double release in the checkpoint reader's streaming handle (D103) StreamHandle.file was a bare var read and written from the reader's worker threads, so two workers could open a descriptor each and the racing assignments released one UncachedFile twice -- the runtime's dangling-reference abort. The discriminator was that every failure was a checkpoint path and every install run was clean, which is InstallFile's versus this cache's unlocked slot. 3 of 4 pair runs failed before, 0 of 5 after, retain-count line zero times. DC-125 is the deterministic regression test. The Python gate is green for the first time in three rounds.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • The head's blocks fan out: 1.7x on the phase, 0.443 tok/s (D102) A whole vocabulary block of rows is the unit, so nothing is regrouped and only the thread changes -- legal where the general matmul's column-range split was not. head 0.434/0.439 -> 0.259/0.255 s/step, digest identical in all four arms; step 2.246 s = 0.443 tok/s. DC-124 is upgraded from unexplained to reproduced-on-the-baseline: the heavy gate-test pair fails when adjacent and passes alone, identically with and without D101/D102.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • The expert misses are fanned across threads: 0.296 -> 0.400 tok/s (D101) preload(experts:shape:) is called once per layer where the router's choices are known, and the adapter fans the misses into the bank as staging. The warm reads do not count as requests but DO count their volume. Bank 512 MB + fan-out: 3.383/3.382 -> 2.503 s/step, mix.read 1.671 -> 0.761, digest identical in all four arms; the bank alone is 4.109, so the win is the concurrency. New row DC-124 records an intermittent abort seen twice and never reproduced, with the isolation attempts the disposition rests on.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • mix.read is read AND unpack, neither is the disk; and code reuse from TinyTitan is now allowed with attribution (D100) SourceTiming surfaced: read 1.298 s/step (38.8%), unpack 1.353 s/step (40.5%), 1.08 GB/step off the device (0.8 GB/s -- not the limit) against 6.5 GB of fp32 materialised per step. The GPU matmul was measured and the disk watchdog stopped it: the head's fp32 array plus an equal MTLBuffer is ~4 GB on a node with ~4.5 GB usable, so DC-120 unblocks it. TinyTitan is Apache-2.0 and taking code is authorised: NOTICE, licence text, modified file marks, and a check_provenance.py change in the same commit as the first file (DC-123).

    @Pummelchen Pummelchen committed Sep 18, 2026
  • The element-wise passes are threaded: 0.242 -> 0.290 tok/s, digest identical (D99) decodeRaw is a map, so threading it cannot change a value: 4.119/4.162 -> 3.448/3.465 s/step, load 1.097 -> 0.356, trace digest identical in all four runs. The matmul was attempted in the same round and REVERTED -- it moved bits and the repo's own reused-buffer comparison caught it, so DC-122 carries the failure rather than the speed-up.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • The expert bank has a lifetime now, and the measurement says capacity is not the lever (D98, DC-119) ExpertBank holds slices for the whole generation with hits, misses, elements read and resident peak reported for the first time. Then five alternated 24-step runs gave 0.0% hits at 0, 512 and 1024 MB with identical 29,192.4M elements read: 773 requests per step of 12.5 MB is a 9.66 GB working set against a 537 MB bank, a reuse distance nine times its capacity. Default is 0, as D89 did for the layer cache; the levers are hiding the reads and cutting their latency.

    @Pummelchen Pummelchen committed Sep 18, 2026
  • Study the sister's decode path and break the 7 tok/s gap into four tasks (D97, DC-118..DC-121) Its decode is a GPU-streaming design: pread streamer with a per-layer slot cache and threaded misses, a GPU-visible residency table, a predictive prefetch ring, and a hard memory clamp at half of physical RAM. Ours has none of them -- and our expert bank is built inside loadLayer and dropped with the layer, which is why our hit rate is 0 by construction rather than by workload.

    @Pummelchen Pummelchen committed Sep 17, 2026
  • Dense models are outside the design: the 4 B measured 0.170 tok/s against the 35 B MoE's 0.230 (D96, DC-116) A dense model has no expert plan to divide and re-reads its whole 3.34 GB payload every token (18.6 GB over four steps, zero hits against the 1 GiB default cache), so it is slower than the sparse 35 B on the same node. The README table gains the model axis; the 4 B artifacts are deleted.

    @Pummelchen Pummelchen committed Sep 17, 2026
  • v1.0.0 published and verified, and one blemish recorded rather than rewritten (D95, DC-115) The release is live with the archive and its checksum; the notes quote b6ebfd1b861a1ac4..., the .sha256 beside the archive matches, shasum -c passes on the downloaded bytes, and all three shipped binaries answer 1.0.0 and report arm64. The archive also carries two TEST bundles because the first version took whatever the build directory held; the selection is now plan-driven and DC-115 says so.

    @Pummelchen Pummelchen committed Sep 17, 2026
  • First release v1.0.0: the identity, the release script and the measured numbers (D95) VERSION is the authority and Version.swift is generated from it; every tool answers --version; the release script runs the gates, does a clean scratch build with a warning scan, asserts lipo -archs on the binaries inside the archive, and refuses notes that do not quote the digest it computed. Measured: 0.23 tok/s prefill and 0.23 tok/s decode on one Mac mini M2, 0.41 tok/s decode across four (1.7x).

    @Pummelchen Pummelchen committed Sep 17, 2026
  • load was one core of eight: the 35 B now runs at 1.70-1.74x over a single node (D94) The int4 dequantiser's row loop is spread across the cores (rows are independent, so it changes which thread computes a value and not how). load 1.33 -> 0.56 s/step single node, and the cluster A/B on one binary gives 1.74x/1.70x with eight threads against 1.66x with one -- four alternated runs, every one bit-identical on all four nodes, on a busy farm, so these are lower bounds. The gate's own >=3x on a quiet farm is still not asserted.

    @Pummelchen Pummelchen committed Sep 17, 2026
  • The head is vocabulary-parallel: 1.13x -> 1.36x (D93) Each node computes its own vocabulary rows with the block decomposition unchanged, then the slices are gathered so every node holds the identical full logits array. Step 4.790 -> 4.021 s, head 1.045 -> 0.26-0.32 s, all four nodes identical. The gap to 1.5x is 0.36 s/step; the exchange counter still does not reconcile with the phases.

    @Pummelchen Pummelchen committed Sep 17, 2026
  • The exchange is 99.6% waiting, and the waiting is the farm (D92) Segments: encode 0.002 s, send 0.007 s, receive 1.11-2.31 s, merge 0.002 s, covering 100.0% of the total. Four interleaved contiguous/round-robin runs (1.06/0.96/0.86/0.88x) falsify the imbalance hypothesis: the slowest cluster step is the run with a peer at load 18.7. A 1.5x figure needs a quiet window.

    @Pummelchen Pummelchen committed Sep 17, 2026
  • The cluster's cost is the exchange (D90); the sister's ANE/GPU split studied (D91) Decode GEMV removes load rather than caching it; ANE prefill is opt-in, fp16 and must verify the device it got; the exchange blocks on peers one at a time, forty times per token. DC-113 added.

    @Pummelchen Pummelchen committed Sep 17, 2026
  • The decoded-layer cache works and on this node it loses (D89) Alternated on one binary: 256 MB neutral, 1 GB 5.46 s/step, 2 GB 5.60 s/step against 5.33 uncached, while load falls as designed. Memory is the constraint, so the default budget is 0 and SHARD_LAYER_CACHE_MB is the knob.

    @Pummelchen Pummelchen committed Sep 17, 2026
  • The cached step's own breakdown: 83% is materialising weights, a third is really I/O (D88) mix.read 33.0%, load 30.5% (CPU dequantisation of constants, not I/O — the payload is read once), head 19.2%. Perfect free sharding of the expert reads is worth 1.09x on the measured step, so the attack is DC-052's decoded-weight budget.

    @Pummelchen Pummelchen committed Sep 17, 2026
  • The first phase breakdown: reads are 64.5% of a forward (D86, D87) mix.read 37.6%, attn.core 23.6%, head 13.5%, load 13.5%. Perfect free sharding of the expert reads is worth only 1.39x, so the 3x target cannot come from the expert plan; the lever is the replicated reads, the attention core and the read path.

    @Pummelchen Pummelchen committed Sep 17, 2026
  • M3's first measurement (0.93x) and the staged-path bug it found (D83, D84) The documented cluster command staged the install under one name and launched with another, so it failed on every peer after copying 21.7 GB to each. Fixed with one shared rule and four tests. The measurement, once fixed: bit-identical on all four nodes, 0.93x, with the exchange budget that explains it.

    @Pummelchen Pummelchen committed Sep 17, 2026