Skip to content

Releases: ashhart/TensorFold

TensorFold 0.3.6.2: EXL3 replies stop on time, and ten community PRs

Choose a tag to compare

@ashhart ashhart released this 28 Sep 16:21

EXL3 replies now stop at the end of their turn, pip installs serve 27B EXL3 packs, and ten community pull requests ship.

  • EXL3 replies stop at the end of the turn. Flash Next and 27B EXL3 packs list <|im_end|> only in generation_config.json, so replies ran past their turn and leaked tool calls and think tags. The CUDA engine now reads that file too. Thanks to @vcruz305 (#69).
  • pip installs serve 27B EXL3 packs. The package was missing the 27B's CUDA sources. A new test checks that every CUDA source ships, and a wheel built from this release carries all 34. Thanks to @taussoe (#66).
  • A quantized KV cache for Flash Next on CUDA. --kv-dtype int8 or int4 holds about 1.7x or 2.6x the default window in the same memory, and drafted replies still equal serial ones. --mtp-confidence sets where MTP chains stop. Thanks to @vcruz305 (#47).
  • Flash Next on CUDA reaches the first token sooner. The head runs on a prompt's final chunk only, and the prompt kernels load at startup, so the first 2k prompt takes 1.26 s instead of 1.79 s. Thanks to @MovieMaker93 (#40).
  • GLM on two Sparks holds 256k tokens with a latent attention cache. Its next-token loss is within 0.001 nats of the per-head cache. Thanks to @taussoe (#54).
  • Mixed-bit Qwen checkpoints (4-bit with some 5- and 6-bit layers, such as oQ4) load on every lane backend, exact, with new row kernels for 2- to 8-bit weights. On an M3 Ultra, oQ4 27B decodes 114-120 tok/s on code and 59-61 on chat, against mlx_lm's 34-36.
  • Concurrent 27B on CUDA: --parallel 16 serves 161.7 tok/s on one Spark in 25.4 GiB, each reply equal to its solo run (#38).
  • Qwen3.6-35B-A3B on CUDA, exact: decode is 1.36-1.49x vLLM with MTP, and prompts are 1.21-1.37x (#45).
  • Gemma 4 drafts (opt-in): --drafter z-lab/gemma-4-26B-A4B-it-DFlash decodes 1.3-2.1x mlx_lm on an M3 Ultra, exact.
  • Tool calls: tool_choice: "required" and a named tool are enforced on both servers (#52). A complete tool call inside an unclosed think block comes back as a tool call (#60).
  • Conversations come back warm on Macs. --spill-gib N writes a conversation pushed out of the prompt cache to disk, up to N GiB, and reads it back when the conversation returns. On a 48 GB budget a 35k-token conversation came back in 0.27 s on an M5 Max instead of 75 s, with the same reply. Off by default; --checkpoint-slots sets how many conversations stay in memory. Thanks to @gilby (#68, #55).
  • Memory: TENSORFOLD_MEMORY_LIMIT_GB raises the budget above the default share, and Flash Next's memory check counts its host-mapped n-gram tables, so it starts on a 128 GB Mac. Thanks to @Chedrian07 (#49, #50).
  • CUDA server fixes from @nood-co1:
    • an abandoned request stops within a round (#57);
    • keys and values stay within the admitted window (#58);
    • a failed admission no longer stops the scheduler (#59).
  • M1-M4: a prompt split into parts now attends exactly as it does in one piece, at every length.

Checked:

  • The CUDA suite on a DGX Spark, built fresh under NVIDIA's container architecture list, with the 27B MLX and EXL3 packs: 883 passed. One admission test failed because it read the Spark's own memory, so it now fakes that as well. The Flash Next EXL3 pack's tests on a second Spark passed too.
  • The Mac suite on an M3 Ultra: 2,183 passed, 0 failed.
  • The Mac suite on an M5 Max: 2,474 passed, 0 failed.

Thanks to @Boscoeuk, @simonmd, @gbgbgbg, @philip-pentatonic and @Deesha08 for the reports (#60, #52, #50, #38, #45, #55).

Install

tensorfold update installs this release. Or:

pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.6.2

TensorFold 0.3.6.1: CUDA builds inside NVIDIA's containers

Choose a tag to compare

@ashhart ashhart released this 28 Sep 09:37

A fix for CUDA builds inside NVIDIA's containers. They set TORCH_CUDA_ARCH_LIST to every architecture back to sm_80, so 0.3.5 to 0.3.6 compiled the kernels' thread-block clusters and FP8 MMA for GPUs without them and stopped with namespace "cooperative_groups" has no member "this_cluster".

  • Every CUDA extension now builds for the GPU that is present, so the container's list adds nothing.
  • A GPU older than compute capability 9.0 is refused with a message that names it.
  • Checked on a DGX Spark under the container's full list with a fresh extension cache: every extension built for sm_121, and the CUDA suite passed (433 tests). The Mac suite passes on an M5 Max (1,779 tests).

Thanks to @ss-cong for the report and the exact errors (#56).

Install

tensorfold update installs this release. Or:

pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.6.1

TensorFold 0.3.6: GLM and Gemma 4 on Macs, bigger models on smaller Macs, EXL3 on NVIDIA GPUs

Choose a tag to compare

@ashhart ashhart released this 28 Sep 09:23

GLM-5.3-Flash and Gemma 4 on Macs, bigger models on smaller Macs, and EXL3 checkpoints on NVIDIA GPUs. Every drafted
reply is still byte-identical to the model's own serial decode, and resumed prompts equal fresh ones.

What's new

  • GLM-5.3-Flash on Macs with 256 GB. It decodes 1.86-2.08x mlx-vlm on an M3 Ultra. Prompts process at or above
    mlx-vlm from 2k to 32k tokens (1.010-1.053x), with a lower peak at long prompts (185 GiB against 213 at 32k).
    Mixed-bit conversions load, and tool calls parse in both servers. Thanks to @chadhurley25075-png (#9, #39),
    @jeidbugs404 (#35) and @kingjamez.
  • Gemma 4 26B-A4B on the lane engine, exact at every width: 9/9 drafted equal to serial, and 1 to 16 concurrent
    streams equal their solo runs. Prompts process at or above mlx_lm from 2k to 64k tokens on an M3 Ultra (1.017-1.117x).
    Thanks to @cshintov (#10).
  • Bigger models on smaller Macs.
    • --ple-on-ssd reads Flash Next's n-gram tables from disk: on a 128 GB budget it peaks at 85.6 GiB and decodes
      at 0.91-1.03x the 256 GB run (#16).
    • --ssd-experts GIB streams routed experts from the checkpoint into a GPU pool of that size. Flash Next fits a
      64 GB Mac (39.5 GiB peak) and GLM a 128 GB one (87.6 GiB), with the resident model's tokens. Decode runs at
      0.31-0.39x resident speed for Flash Next and 0.13-0.17x for GLM. Install with pip install "tensorfold[ssd]"
      (#17). Both were measured on an M3 Ultra under emulated budgets.
  • EXL3 checkpoints on NVIDIA GPUs (experimental): turboderp's Qwen3.8-27B and Flash Next packs, exact on the lanes
    (9/9, resume 6/6, four concurrent streams equal solo). On one DGX Spark against vLLM with MTP=3, the 27B decodes
    3.56 / 1.74 / 2.50 / 1.62x and Flash Next 1.91 / 1.79 / 1.90 / 1.84x (code sampled / chat sampled / code greedy /
    chat greedy). Flash Next's pack admits its full 262,144-token window on one Spark. Prompt processing is about half
    the MLX checkpoints' speed for now. Thanks to @vcruz305 (#42).
  • Faster prompts. Flash Next sizes its prompt chunks to the memory it has: 1.20-1.38x 0.3.5.1 on an M3 Ultra,
    and level with oMLX from 8k to 64k. Nemotron takes up to 8,192 tokens a chunk on M5 GPUs, 12-15% faster at 8k and
    32k. The weights stay wired in memory while a server runs.
  • tensorfold update shows what's new when it finishes, from CHANGELOG.md, and the first run
    of a new version links it.
  • MLX 0.32.2 or newer is required on Macs.

Fixed

  • GLM-5.3-Flash on two Sparks answered "!" past about 2,000 prompt tokens: a prompt chunk's rows shared their
    sparse-attention partials. EXL3 GLM prompts past 128 tokens failed too (#53, reported by @taussoe).
  • A reply that isn't a tool call comes back as content instead of an HTTP 500 (#51, reported by @simonmd).
  • After a long prompt the server gives back 7.5-7.9 GiB at rest, clearing MLX's buffer cache only when nothing is
    decoding (#44, @kingjamez).

Checks on this release

  • The test suite passes: 1,778 tests on an M5 Max and 1,482 on an M3 Ultra, with no failures. The four known M1 to
    M4 cases below are marked as known.
  • Every landed change kept drafted == serial and resumed == fresh on the machines it touches, with token hashes
    equal to the tree before it wherever the output shouldn't change.

Where 0.3.6 is still short

  • EXL3 prompt processing runs at about half the MLX checkpoints' speed. An FP8 prompt path is next.
  • The 3x decode floor isn't met on every model: Gemma 4 is at 1.10-1.45x mlx_lm, GLM on Macs at 1.86-2.08x mlx-vlm,
    and the EXL3 cells are as listed above.
  • Flash Next's prompt processing at 32k and 64k is 1-1.4% under oMLX on an M3 Ultra.
  • On M1 to M4 with MLX 0.32.2, the 27B's prompt attention at 8,192 keys doesn't match one stock MLX call bit for bit
    when a chunk splits with a short tail. 0.3.5.1 behaves the same. Drafted and serial output, and resumed and fresh
    prompts, still agree.
  • Replies to prompts longer than one prompt chunk can differ between machines with different memory, because the
    chunk size follows the memory budget.
  • Next in 0.3.6.1: GLM's latent cache for 256k windows (#54), tool_choice: required (#52), explicit memory budgets
    (#49, #50), the 27B at 16 concurrent streams on CUDA and the Qwen3.6-35B-A3B family (#38, #45), quantized KV and
    Flash Next prefill on CUDA (#47, #40), Ternary Bonsai (#18), more quantized checkpoints, image input, and
    DeepSeek-V4-Flash (#14).

Install

tensorfold update installs this release. Or:

pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.6

TensorFold 0.3.5.1

Choose a tag to compare

@ashhart ashhart released this 28 Sep 02:49

A fix for M1 and M2 Macs: Qwen3.8-27B wouldn't load on them. 0.3.4.1 stopped with "Thread group size (1024) is greater than the maximum allowed threads per threadgroup (704)", and 0.3.5 with 512 over 448.

Metal gives each compiled kernel its own limit on threads per threadgroup, and on M1 and M2 that limit falls as the kernel uses more registers. M3 and later give every kernel 1024, which is why our test machines never hit it.

  • On M1 and M2 the 4-bit row matmul checks each of its kernels the first time it runs and uses fewer simdgroups where 16 don't fit. The sums run in the same order, so drafted output still equals serial output.
  • Every other kernel above 256 threads now declares its size to the Metal compiler, which makes M1 and M2 fit it: the norms, sampling and top-k, Nemotron's norms and row matmul, and Flash Next's larger kernels.
  • On M3, M4 and M5 nothing changes: the compiled kernels are 0.3.5's machine code. Paired against 0.3.5 on an M3 Ultra, every token hash matches, and decode and cold prefill from 2k to 32k tokens stay level with 0.3.5 (0.99-1.02x).

Thanks to @hichaiuse, @simonmd, @gcarusso, @tonydehnke, @Cyb3r-Monk and @tinyapps for the reports, the repro and the numbers that found it.

Install

tensorfold update installs this release. Or:

pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.5.1

TensorFold 0.3.5: concurrent requests, resume at message starts, memory that fits

Choose a tag to compare

@ashhart ashhart released this 27 Sep 23:13

Concurrent requests now share each verification round, follow-up turns resume from the start of their newest messages, and the server keeps its whole memory footprint inside what the Mac can hold. Every reply is still byte-identical to the model's own serial decode, and each concurrent stream equals its solo run.

What's new

  • Concurrency. --parallel auto is on by default: requests share one verification round, so aggregate throughput rises with the number of streams. On Metal and CUDA, every stream's reply equals the same request served alone.
  • Resume at message starts. Prompts are chunked where replies begin, so a follow-up turn prefills only its latest reply and new messages, with output identical to a fresh prompt. On a 12-turn agent session with the 27B, first-token time summed to 14.9 s instead of 31.9 s. --prefill-grid is gone.
  • Memory. The whole process stays inside 70% of RAM, an omitted --context defaults to the window the machine can hold (printed at startup), and a prompt past it gets a clear 400 instead of pushing macOS into swap. The README lists tested windows per memory size.
  • Flash Next prefill. @quigles1977's prefill path (#29) with sparse prompt attention: cold prompts 1.1-1.4x faster than 0.3.4.1 on an M3 Ultra, and no pause before the first decoded token after a long prompt.
  • 2- to 8-bit weights. @jasontitus's 3- and 2-bit lane kernels (#34), extended to 5-, 6- and 8-bit, so mixed-precision 27B checkpoints decode fully on the lanes. Checkpoints the lanes can't read are refused from config.json before any download.
  • CUDA. FP8 prefill and shared expert kernels (prefill 2.5-3x faster than 0.3.4.1 for the 27B and about 9x for Flash Next on one DGX Spark), concurrent streams for the 27B and Flash Next, admission from available memory before loading, and the affordable native context by default.
  • API. Raw /v1/completions prompts (no chat template, no think block), ignore_eos and stop strings, request reasoning_effort and typed tool arguments (@chris247474, #28), developer messages as system instructions, parallel_tool_calls: false, tools with null parameters, and cancellation when a client disconnects. An explicit max_tokens that doesn't fit the context now gets a 400; 0.3.4.1 capped it.

Fixed

#19, #20, #21, #23, #25, #26, #27, #31, #32, #36, #46.

Measured on this release

Paired against 0.3.4.1 on the same machine, cold prompts, thinking off; decode cells are 64-token replies (code sampled / chat sampled / code greedy / chat greedy).

Model Machine Exact Decode against 0.3.4.1 Prefill against 0.3.4.1
Qwen3.8 Flash Next M3 Ultra 22/22, resume 2/2, 16 streams equal solo 1.001 / 0.999 / 0.997 / 1.014 1.09x at 2k, 1.12x at 8k, 1.36x at 32k
Nemotron 3.5 Lightning M5 Max 22/22, resume 2/2, 16 streams equal solo 0.984 / 0.999 / 0.969 / 0.999 1.05x at 2k, 1.00x at 8k, 1.01x at 32k
Qwen3.8-27B M5 Max 22/22, resume 2/2, 16 streams equal solo about 0.97 a round (see below) 1.00x at 2k, 1.03x at 8k, 1.01x at 16k, 1.03x at 32k
Qwen3.8-27B M3 Ultra 22/22, resume 2/2, 16 streams equal solo 1.196 / 1.171 / 1.181 / 1.200 0.99x at 2k, 8k and 32k
Qwen3.8-27B DGX Spark 9/9 1.12 / 1.13 / 1.12 / 1.13 2.5-3x
Qwen3.8 Flash Next DGX Spark 9/9 1.18 / 1.08 / 1.03 / 1.23 about 9x
GLM-5.3-Flash two DGX Sparks 3/3 1.088 / 1.012 / 0.973 / 1.011

Where 0.3.5 is still short:

  • Flash Next prompt processing is 0.80-0.84x of a stock oMLX server on the same M3 Ultra, and Nemotron's is 0.82-0.85x of mlx_lm at 2k-8k on an M5 Max (level at 32k). 0.3.4.1 had the same gaps; closing them is next.
  • GLM on two Sparks decodes code greedily 2.7% slower than 0.3.4.1; the automatic drafter choice now re-probes a drafter it left after three rounds, and the rest is the next fix. Nemotron's code-sampled cell reads 1.6% under 0.3.4.1 at the median.
  • Resuming at message starts costs cold prefill on prompts with many messages: 7.8% on a 40-message 18.7k session for the 27B, 4.3% for Nemotron, 2.9% for Flash Next. Single-message prompts are level.
  • The 27B on M5 decodes about 3% slower than 0.3.4.1 (each round ~1.3 ms dearer, same tokens) in the new engine path; a bounds check that cost another ~8% was removed before release. The rest follows in 0.3.5.1.
  • Before M5, Flash Next decodes with 0.3.4.1's per-row kernels, so a single stream matches 0.3.4.1 exactly; faster kernels at the same bits are next.

Contributors

GLM-5.3-Flash on Mac (#9, #35, #39), Gemma 4 (#10) and the CUDA prefill and EXL3 work (#40, #42) follow in the next releases.

Install

tensorfold update installs this release. Or:

pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.5

TensorFold 0.3.4.1: prompt processing back to full speed

Choose a tag to compare

@ashhart ashhart released this 27 Sep 08:47

A fix for 0.3.4: prompt processing is back to MLX's speed on every model, and a resumed conversation still matches the same conversation fed fresh, byte for byte.

What was wrong

0.3.4 processed Nemotron's and Flash Next's prompts through their row-exact decode kernels, so that a conversation resumed from a cached prefix got exactly the bits of a fresh one. That made prompt processing several times slower than MLX's own: about 550 tok/s for Nemotron on an M5 Max, where mlx_lm runs about 2,500-3,000, and about 300 for Flash Next on an M3 Ultra, where MLX runs 760.

What changed

  • Every model prefills through MLX's own forward, in chunks on a fixed 2,048-token grid from the start of the prompt. Prompt caches are kept only at grid points, and a reply is prefilled again from the last grid point on the next turn. A chunk's bits then never depend on where a conversation was resumed, so resumed conversations still equal fresh ones.
  • Flash Next queues at most two layers ahead of the GPU while it prefills, so its peak memory stays bounded at any context length.
  • The server no longer holds the first request while it warms saved system blocks at startup; the warm runs in the background.
  • Prompt caches saved by 0.3.4 are not reused, because their key now names the prefill mode, so the first prompt after upgrading is prefilled from scratch.

Measured on this tree

Cold prompts, tok/s, through tensorfold serve:

Model Machine Prompt 0.3.4 0.3.4.1 MLX
Nemotron 3.5 Lightning M5 Max 8k ~550 3,400 2,540 (mlx_lm, same session)
32k ~540 2,456 2,448 (mlx_lm, same session)
Qwen3.8-27B M5 Max 2k 542 867 812 (mlx_lm)
32k 510 560 562 (mlx_lm)
64k 441 465 467 (mlx_lm)
Qwen3.8 Flash Next M3 Ultra 2k 305 934
16k ~290 820 759 (MLX, at 22.8k)
64k ~280 550
196k ~255 315

Flash Next's peak memory stays within 20 GB of its weights up to a 196k-token prompt (134 GB peak on the M3 Ultra). Its prompt speed still falls with context length; a sparse prefill kernel that reads only each row's selected keys is next.

Checks on this tree:

  • Every drafted reply equals the same request sent with "draft": false, 9 of 9 on each model.
  • Resumed multi-turn conversations equal the same conversations fed fresh, including a three-turn Flash Next chat resumed at nine grid points of a 19k-token first message, with thinking on and off.
  • Decode speed is unchanged: Nemotron on an M5 Max within noise of 0.3.4 over 15 seeds a cell, and Flash Next on an M3 Ultra at 135.5 / 111.1 / 142.3 / 122.2 against 134.1 / 111.8 / 137.7 / 117.0 (code sampled / chat sampled / code greedy / chat greedy).
  • The test suite passes (370 tests).

Install

tensorfold update installs this release. Or:

pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.4.1

TensorFold 0.3.4 (work in progress)

Pre-release

Choose a tag to compare

@ashhart ashhart released this 26 Sep 22:47
2f8e514

Work in progress. 0.3.4 publishes everything built since 0.3.3 as a working snapshot. It's marked as a pre-release, so tensorfold update keeps offering 0.3.3; install it on purpose with the command at the end. Every number below was measured on this tree. Every drafted reply was checked byte for byte against the same request sent with "draft": false.

Every model on the lanes

TensorFold drafts tokens and verifies them in batched "lanes": several drafted rows go through one forward. Before this release, only Qwen3.8-27B ran on the lane engine; Nemotron and Flash Next decoded on a separate serial engine with narrow drafting windows. Now:

  • One engine for everything. Nemotron 3.5 Lightning and Qwen3.8 Flash Next run on the lane engine through family rounds, and the serial engine is gone. Every drafted round verifies at least 2 rows, and the server reports min_rows and a token-id hash (token_sha) for every reply.
  • Exact rows on every M chip. Nemotron has row-exact kernels on Macs without the M5's tensor units: windows of 2 to 16 rows reproduce one-row decoding there. So on M1 to M4 it drafts instead of decoding one token at a time. On the M5, every verify matmul (dense and experts) is now TensorFold's own, so exactness no longer depends on the MLX version.
  • Resumed prompts equal fresh ones. Nemotron and Flash Next now process prompts through the same row-exact kernels as decoding, so a resumed conversation reproduces the same conversation served fresh. Before, MLX's chunked prefill could make them differ at near-ties.
  • Cheaper drafts. Nemotron's MTP head drafts over a 32,768-id vocabulary and Flash Next's over its 79,592-id list, so a draft step reads a quarter of the head. Both lists are built from public text only, and their corpora are recorded next to them (Nemotron's: CPython's standard library plus this package's own source and docs). Flash Next's chained drafts stay on the GPU.

Qwen3.8-27B on M1 to M4: 1.9 to 4x serial

0.3.3 made drafting exact on Macs without tensor units. 0.3.4 makes it fast:

  • A new lane decoder for these chips (kernels/qwen/dense/v1/row_forward.py): fused glue, stacked projections and the lane recurrence.
  • A new 4-bit matmul on the simdgroup matrix units every Apple GPU has (kernels/qwen/dense/v1/simd_qmm.py). It dequantizes each 8x8 weight tile once and multiplies up to 16 rows, and a row's bits never depend on how many rows ride with it. On an M3 Ultra a window of 2 to 8 rows costs about 32 ms, against 25 ms for one row.
  • Serial decoding, windows and prompts all go through it, so it defines the reference the drafted rounds reproduce.
  • Prompts resume only from a 2,048-token grid (TF_ROW_PREFILL=aligned, the default), so a resumed conversation equals a fresh one at MLX's prompt speed. This fixes a 0.3.3 gap: on M1 to M4 a resumed conversation could differ from a fresh one at near-ties (13 of 25 follow-ups in a stress test).
  • Tool-calling requests on these chips now draft only as many tokens as pay. Before, they always drafted 7.

M3 Ultra, MLX 0.32.0, 64-token replies, thinking off (code sampled / chat sampled / code greedy / chat greedy):

tok/s
MLX serial 38.2 / 38.2 / 39.3 / 39.3
0.3.3 63.9 / 47.0 / 64.4 / 51.6
0.3.4 141.3 / 73.9 / 158.4 / 74.2

CUDA: lanes on every default path

An audit of every decode path in the CUDA engines found five default configurations that could decode one token a forward. All five are fixed:

  • Flash Next always verifies its first MTP draft, so every round is at least 2 rows.
  • The 27B refuses to start without its DFlash2 draft model, and Flash Next without its MTP head, each naming the fix (--no-drafts still serves the serial reference). Before, both served mostly one token a round.
  • GLM-5.3-Flash with --mtp-drafts 0, or a checkpoint without its MTP head, drafts with DFlash2. Before, the first left the loaded draft model idle and the second crashed.

The Spark speeds are unchanged within 1%. On one DGX Spark, the 27B runs 50.0 / 45.5 / 49.3 / 46.0 tok/s and Flash Next 64.6 / 56.9 / 72.5 / 59.2, with 9/9 drafted equal to serial and prefix reuse exact.

Measured on this tree

Same cells: bench_openai, 64-token replies, thinking off.

Model Machine Code sampled Chat sampled Code greedy Chat greedy Against
Nemotron 3.5 Lightning M5 Max 288.0 222.9 291.8 242.6 mlx_lm's own server 172.6 / 172.8 / 178.6 / 175.9
Nemotron 3.5 Lightning M3 Ultra 295.1 277.0 332.8 278.5 0.3.3 served it serially there (~220)
Qwen3.8 Flash Next M3 Ultra 127.4 111.1 137.7 117.8 0.3.3: 122.7 / 109.8 / 141.2 / 115.3
Qwen3.8-27B + DFlash2 M5 Max 153.2 69.8 154.2 72.8 unchanged from 0.3.3
Qwen3.8-27B + DFlash2 M3 Ultra 141.3 73.9 158.4 74.2 MLX serial 38.2 / 38.2 / 39.3 / 39.3

Every row passed 9 of 9 drafted replies equal to "draft": false, and resumed chats equal to fresh ones.

Known issues

  • Some cells are still under twice the standard. That is our floor. The 27B's chat cells on M1 to M4 are at 1.9x, Nemotron's chat cells at 1.3-1.4x mlx_lm, Flash Next at 1.4-1.7x MLX serial on an M3 Ultra, and several CUDA cells.
  • Cold long prompts are slower for Nemotron and Flash Next. Exact prompt processing costs speed there: Flash Next about 281 against 759 tok/s on an M3 Ultra, Nemotron about 650 tok/s on an M5 Max. Prompts resumed from a cached prefix are not affected.
  • Concurrent requests are served one at a time. Several requests drafting together in one forward is in progress. The first measurement is 4 concurrent 27B requests at 177 tok/s combined on an M5 Max, all 4 exact, but it isn't in this release.
  • Draft trees on M1 to M4 are opt-in (TF_ROW_ATTENTION=1); they're exact but slower than chains today.

Next experiment: MTP heads and drafters

The lanes are cheap now. On an M5 Max the 27B verifies 16 rows in about 36 ms and 128 rows in 145, and 128 lanes can carry about 620 tokens a second. What limits a single request is how many of those lanes hold right guesses, which comes down to the drafter:

  • The 27B: chat takes 3.5-3.95 tokens a round, where DFlash2's own top candidates would allow 5.6-6.1.
  • Nemotron: on chat, its MTP head's first 1, 2, 3 and 4 chained drafts are all right 74%, 48%, 31% and 16% of the time.
  • Flash Next on a Spark: perfect use of today's MTP drafts still leaves the code cells short of twice vLLM.

Tuning trees, draft counts, copies and scoring on today's drafters moves the M5 by no more than ±2%. So the next experiment is the drafters themselves:

  1. Fine-tune each model's own MTP head on the model's own output, unrolled several drafts deep the way the engine chains them.
  2. A TensorFold drafter that returns a whole calibrated tree in one small forward, trained on what the engine actually accepts.

Both keep output byte-identical, since drafts only change speed. The goal is every cell at least twice the standard, on every Mac and on DGX Spark.

Install

pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.4

TensorFold 0.3.3

Choose a tag to compare

@ashhart ashhart released this 26 Sep 21:11
b08e014

All Apple Silicon chips now support lane batching. Qwen3.8-27B verifies its drafted tokens together in one forward on every M1 to M5 GPU, with output byte-identical to serial decoding.

Before 0.3.3, Macs without the M5's tensor units served the 27B at serial speed: the DFlash2 drafter stopped after its first round. Now:

  • M1 to M4 GPUs verify drafted windows of 2 to 8 rows through a new row-exact 4-bit matvec (kernels/qwen/dense/v1/row_qmv.py). A row gets the same bits whatever the window size, so drafted output still equals serial decoding.
  • At load the engine checks which window widths reproduce one-row decoding on your Mac and times them. Each round then drafts as many tokens as pay at the request's recent acceptance.
  • M5 GPUs keep the lane kernels, and their speed is unchanged.

Measured on an M3 Ultra (MLX 0.32.0), 64-token replies, median of seeds 1234-1238, through tools/bench_openai.py:

Code, sampled Chat, sampled Code, greedy Chat, greedy
Serial (--no-drafts) 38.2 38.2 39.3 39.3
0.3.3, DFlash2 drafts 63.9 47.0 64.4 51.6

Every drafted reply equaled the same request sent with "draft": false (9 of 9), and resumed conversations gave identical replies. The 27B recipe has the details.

Also fixed: streamed /v1/completions on the Mac server sent reasoning as an object in "text". Text completions now stream plain text, as their non-streamed reply does.

Update from 0.3.2 with tensorfold update, or from an earlier release with:

pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.3

TensorFold 0.3.2

Choose a tag to compare

@ashhart ashhart released this 26 Sep 20:12
a83ed1e

TensorFold 0.3.2 adds two things: GLM-5.3-Flash on Mia-AiLab's EXL3 weights, as an experiment, and tensorfold update.

GLM-5.3-Flash on EXL3 (experimental). The CUDA engine now reads Mia-AiLab/GLM-5.3-Flash-EXL3-TR3-4bpw, the checkpoint vLLM serves in Mia-AiLab's DGX Spark recipe, on two Sparks. On those same weights, through the same OpenAI client (decode tok/s, one stream, 64-token replies, median of seeds 1234-1238):

Code, sampled Chat, sampled Code, greedy Chat, greedy
vLLM, Mia-AiLab's recipe, MTP=3 24.5 24.3 32.2 24.7
TensorFold 0.3.2 36.4 29.7 43.8 32.9

Replies are byte-identical to TensorFold's own serial decoding, as on every other model. The MLX checkpoint (Vontra/GLM-5.3-Flash-MLX-4bit-MTP) is still the faster way to run GLM here, because only the EXL3 checkpoint's routed experts are 4-bit: each Spark reads 10.7 GB a token from it against 5.0. The recipe has the kernels, the checks and what was not checked. The EXL3 format is ExLlamaV3's (MIT).

tensorfold update. tensorfold update installs the newest release. serve also checks GitHub for a newer release once a day in the background, and prints a line if it finds one. It never delays the server. Turn it off with --no-update-check or TENSORFOLD_NO_UPDATE_CHECK=1. From 0.3.1 or earlier, update once by hand:

pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.2

TensorFold 0.3.1

Choose a tag to compare

@ashhart ashhart released this 26 Sep 18:27
d7470ed

TensorFold 0.3.1 says plainly when a checkpoint has no recipe, before anything downloads.

  • tensorfold info MODEL shows how a checkpoint stores its weights (MLX 4-bit, groups of 64, exl3 (4-bit), ...) and which backends read them.
  • serve and pull refuse a model type with no family, or weights no engine reads yet (EXL3, NVFP4, GPTQ, AWQ), and point to the recipe book and the runbook.
  • A checkpoint that isn't one TensorFold is tested with still runs when its format matches, with a note that its speed and quality are unmeasured.

Upgrade: pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.1.