Releases: ashhart/TensorFold
Release list
TensorFold 0.3.6.2: EXL3 replies stop on time, and ten community PRs
EXL3 replies now stop at the end of their turn, pip installs serve 27B EXL3 packs, and ten community pull requests ship.
- EXL3 replies stop at the end of the turn. Flash Next and 27B EXL3 packs list
<|im_end|>only ingeneration_config.json, so replies ran past their turn and leaked tool calls and think tags. The CUDA engine now reads that file too. Thanks to @vcruz305 (#69). - pip installs serve 27B EXL3 packs. The package was missing the 27B's CUDA sources. A new test checks that every CUDA source ships, and a wheel built from this release carries all 34. Thanks to @taussoe (#66).
- A quantized KV cache for Flash Next on CUDA.
--kv-dtype int8orint4holds about 1.7x or 2.6x the default window in the same memory, and drafted replies still equal serial ones.--mtp-confidencesets where MTP chains stop. Thanks to @vcruz305 (#47). - Flash Next on CUDA reaches the first token sooner. The head runs on a prompt's final chunk only, and the prompt kernels load at startup, so the first 2k prompt takes 1.26 s instead of 1.79 s. Thanks to @MovieMaker93 (#40).
- GLM on two Sparks holds 256k tokens with a latent attention cache. Its next-token loss is within 0.001 nats of the per-head cache. Thanks to @taussoe (#54).
- Mixed-bit Qwen checkpoints (4-bit with some 5- and 6-bit layers, such as oQ4) load on every lane backend, exact, with new row kernels for 2- to 8-bit weights. On an M3 Ultra, oQ4 27B decodes 114-120 tok/s on code and 59-61 on chat, against mlx_lm's 34-36.
- Concurrent 27B on CUDA:
--parallel 16serves 161.7 tok/s on one Spark in 25.4 GiB, each reply equal to its solo run (#38). - Qwen3.6-35B-A3B on CUDA, exact: decode is 1.36-1.49x vLLM with MTP, and prompts are 1.21-1.37x (#45).
- Gemma 4 drafts (opt-in):
--drafter z-lab/gemma-4-26B-A4B-it-DFlashdecodes 1.3-2.1x mlx_lm on an M3 Ultra, exact. - Tool calls:
tool_choice: "required"and a named tool are enforced on both servers (#52). A complete tool call inside an unclosed think block comes back as a tool call (#60). - Conversations come back warm on Macs.
--spill-gib Nwrites a conversation pushed out of the prompt cache to disk, up to N GiB, and reads it back when the conversation returns. On a 48 GB budget a 35k-token conversation came back in 0.27 s on an M5 Max instead of 75 s, with the same reply. Off by default;--checkpoint-slotssets how many conversations stay in memory. Thanks to @gilby (#68, #55). - Memory:
TENSORFOLD_MEMORY_LIMIT_GBraises the budget above the default share, and Flash Next's memory check counts its host-mapped n-gram tables, so it starts on a 128 GB Mac. Thanks to @Chedrian07 (#49, #50). - CUDA server fixes from @nood-co1:
- M1-M4: a prompt split into parts now attends exactly as it does in one piece, at every length.
Checked:
- The CUDA suite on a DGX Spark, built fresh under NVIDIA's container architecture list, with the 27B MLX and EXL3 packs: 883 passed. One admission test failed because it read the Spark's own memory, so it now fakes that as well. The Flash Next EXL3 pack's tests on a second Spark passed too.
- The Mac suite on an M3 Ultra: 2,183 passed, 0 failed.
- The Mac suite on an M5 Max: 2,474 passed, 0 failed.
Thanks to @Boscoeuk, @simonmd, @gbgbgbg, @philip-pentatonic and @Deesha08 for the reports (#60, #52, #50, #38, #45, #55).
Install
tensorfold update installs this release. Or:
pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.6.2TensorFold 0.3.6.1: CUDA builds inside NVIDIA's containers
A fix for CUDA builds inside NVIDIA's containers. They set TORCH_CUDA_ARCH_LIST to every architecture back to sm_80, so 0.3.5 to 0.3.6 compiled the kernels' thread-block clusters and FP8 MMA for GPUs without them and stopped with namespace "cooperative_groups" has no member "this_cluster".
- Every CUDA extension now builds for the GPU that is present, so the container's list adds nothing.
- A GPU older than compute capability 9.0 is refused with a message that names it.
- Checked on a DGX Spark under the container's full list with a fresh extension cache: every extension built for sm_121, and the CUDA suite passed (433 tests). The Mac suite passes on an M5 Max (1,779 tests).
Thanks to @ss-cong for the report and the exact errors (#56).
Install
tensorfold update installs this release. Or:
pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.6.1TensorFold 0.3.6: GLM and Gemma 4 on Macs, bigger models on smaller Macs, EXL3 on NVIDIA GPUs
GLM-5.3-Flash and Gemma 4 on Macs, bigger models on smaller Macs, and EXL3 checkpoints on NVIDIA GPUs. Every drafted
reply is still byte-identical to the model's own serial decode, and resumed prompts equal fresh ones.
What's new
- GLM-5.3-Flash on Macs with 256 GB. It decodes 1.86-2.08x mlx-vlm on an M3 Ultra. Prompts process at or above
mlx-vlm from 2k to 32k tokens (1.010-1.053x), with a lower peak at long prompts (185 GiB against 213 at 32k).
Mixed-bit conversions load, and tool calls parse in both servers. Thanks to @chadhurley25075-png (#9, #39),
@jeidbugs404 (#35) and @kingjamez. - Gemma 4 26B-A4B on the lane engine, exact at every width: 9/9 drafted equal to serial, and 1 to 16 concurrent
streams equal their solo runs. Prompts process at or above mlx_lm from 2k to 64k tokens on an M3 Ultra (1.017-1.117x).
Thanks to @cshintov (#10). - Bigger models on smaller Macs.
--ple-on-ssdreads Flash Next's n-gram tables from disk: on a 128 GB budget it peaks at 85.6 GiB and decodes
at 0.91-1.03x the 256 GB run (#16).--ssd-experts GIBstreams routed experts from the checkpoint into a GPU pool of that size. Flash Next fits a
64 GB Mac (39.5 GiB peak) and GLM a 128 GB one (87.6 GiB), with the resident model's tokens. Decode runs at
0.31-0.39x resident speed for Flash Next and 0.13-0.17x for GLM. Install withpip install "tensorfold[ssd]"
(#17). Both were measured on an M3 Ultra under emulated budgets.
- EXL3 checkpoints on NVIDIA GPUs (experimental): turboderp's Qwen3.8-27B and Flash Next packs, exact on the lanes
(9/9, resume 6/6, four concurrent streams equal solo). On one DGX Spark against vLLM with MTP=3, the 27B decodes
3.56 / 1.74 / 2.50 / 1.62x and Flash Next 1.91 / 1.79 / 1.90 / 1.84x (code sampled / chat sampled / code greedy /
chat greedy). Flash Next's pack admits its full 262,144-token window on one Spark. Prompt processing is about half
the MLX checkpoints' speed for now. Thanks to @vcruz305 (#42). - Faster prompts. Flash Next sizes its prompt chunks to the memory it has: 1.20-1.38x 0.3.5.1 on an M3 Ultra,
and level with oMLX from 8k to 64k. Nemotron takes up to 8,192 tokens a chunk on M5 GPUs, 12-15% faster at 8k and
32k. The weights stay wired in memory while a server runs. tensorfold updateshows what's new when it finishes, from CHANGELOG.md, and the first run
of a new version links it.- MLX 0.32.2 or newer is required on Macs.
Fixed
- GLM-5.3-Flash on two Sparks answered "!" past about 2,000 prompt tokens: a prompt chunk's rows shared their
sparse-attention partials. EXL3 GLM prompts past 128 tokens failed too (#53, reported by @taussoe). - A reply that isn't a tool call comes back as content instead of an HTTP 500 (#51, reported by @simonmd).
- After a long prompt the server gives back 7.5-7.9 GiB at rest, clearing MLX's buffer cache only when nothing is
decoding (#44, @kingjamez).
Checks on this release
- The test suite passes: 1,778 tests on an M5 Max and 1,482 on an M3 Ultra, with no failures. The four known M1 to
M4 cases below are marked as known. - Every landed change kept drafted == serial and resumed == fresh on the machines it touches, with token hashes
equal to the tree before it wherever the output shouldn't change.
Where 0.3.6 is still short
- EXL3 prompt processing runs at about half the MLX checkpoints' speed. An FP8 prompt path is next.
- The 3x decode floor isn't met on every model: Gemma 4 is at 1.10-1.45x mlx_lm, GLM on Macs at 1.86-2.08x mlx-vlm,
and the EXL3 cells are as listed above. - Flash Next's prompt processing at 32k and 64k is 1-1.4% under oMLX on an M3 Ultra.
- On M1 to M4 with MLX 0.32.2, the 27B's prompt attention at 8,192 keys doesn't match one stock MLX call bit for bit
when a chunk splits with a short tail. 0.3.5.1 behaves the same. Drafted and serial output, and resumed and fresh
prompts, still agree. - Replies to prompts longer than one prompt chunk can differ between machines with different memory, because the
chunk size follows the memory budget. - Next in 0.3.6.1: GLM's latent cache for 256k windows (#54),
tool_choice: required(#52), explicit memory budgets
(#49, #50), the 27B at 16 concurrent streams on CUDA and the Qwen3.6-35B-A3B family (#38, #45), quantized KV and
Flash Next prefill on CUDA (#47, #40), Ternary Bonsai (#18), more quantized checkpoints, image input, and
DeepSeek-V4-Flash (#14).
Install
tensorfold update installs this release. Or:
pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.6TensorFold 0.3.5.1
A fix for M1 and M2 Macs: Qwen3.8-27B wouldn't load on them. 0.3.4.1 stopped with "Thread group size (1024) is greater than the maximum allowed threads per threadgroup (704)", and 0.3.5 with 512 over 448.
Metal gives each compiled kernel its own limit on threads per threadgroup, and on M1 and M2 that limit falls as the kernel uses more registers. M3 and later give every kernel 1024, which is why our test machines never hit it.
- On M1 and M2 the 4-bit row matmul checks each of its kernels the first time it runs and uses fewer simdgroups where 16 don't fit. The sums run in the same order, so drafted output still equals serial output.
- Every other kernel above 256 threads now declares its size to the Metal compiler, which makes M1 and M2 fit it: the norms, sampling and top-k, Nemotron's norms and row matmul, and Flash Next's larger kernels.
- On M3, M4 and M5 nothing changes: the compiled kernels are 0.3.5's machine code. Paired against 0.3.5 on an M3 Ultra, every token hash matches, and decode and cold prefill from 2k to 32k tokens stay level with 0.3.5 (0.99-1.02x).
Thanks to @hichaiuse, @simonmd, @gcarusso, @tonydehnke, @Cyb3r-Monk and @tinyapps for the reports, the repro and the numbers that found it.
Install
tensorfold update installs this release. Or:
pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.5.1TensorFold 0.3.5: concurrent requests, resume at message starts, memory that fits
Concurrent requests now share each verification round, follow-up turns resume from the start of their newest messages, and the server keeps its whole memory footprint inside what the Mac can hold. Every reply is still byte-identical to the model's own serial decode, and each concurrent stream equals its solo run.
What's new
- Concurrency.
--parallel autois on by default: requests share one verification round, so aggregate throughput rises with the number of streams. On Metal and CUDA, every stream's reply equals the same request served alone. - Resume at message starts. Prompts are chunked where replies begin, so a follow-up turn prefills only its latest reply and new messages, with output identical to a fresh prompt. On a 12-turn agent session with the 27B, first-token time summed to 14.9 s instead of 31.9 s.
--prefill-gridis gone. - Memory. The whole process stays inside 70% of RAM, an omitted
--contextdefaults to the window the machine can hold (printed at startup), and a prompt past it gets a clear 400 instead of pushing macOS into swap. The README lists tested windows per memory size. - Flash Next prefill. @quigles1977's prefill path (#29) with sparse prompt attention: cold prompts 1.1-1.4x faster than 0.3.4.1 on an M3 Ultra, and no pause before the first decoded token after a long prompt.
- 2- to 8-bit weights. @jasontitus's 3- and 2-bit lane kernels (#34), extended to 5-, 6- and 8-bit, so mixed-precision 27B checkpoints decode fully on the lanes. Checkpoints the lanes can't read are refused from
config.jsonbefore any download. - CUDA. FP8 prefill and shared expert kernels (prefill 2.5-3x faster than 0.3.4.1 for the 27B and about 9x for Flash Next on one DGX Spark), concurrent streams for the 27B and Flash Next, admission from available memory before loading, and the affordable native context by default.
- API. Raw
/v1/completionsprompts (no chat template, no think block),ignore_eosandstopstrings, requestreasoning_effortand typed tool arguments (@chris247474, #28),developermessages as system instructions,parallel_tool_calls: false, tools with null parameters, and cancellation when a client disconnects. An explicitmax_tokensthat doesn't fit the context now gets a 400; 0.3.4.1 capped it.
Fixed
#19, #20, #21, #23, #25, #26, #27, #31, #32, #36, #46.
Measured on this release
Paired against 0.3.4.1 on the same machine, cold prompts, thinking off; decode cells are 64-token replies (code sampled / chat sampled / code greedy / chat greedy).
| Model | Machine | Exact | Decode against 0.3.4.1 | Prefill against 0.3.4.1 |
|---|---|---|---|---|
| Qwen3.8 Flash Next | M3 Ultra | 22/22, resume 2/2, 16 streams equal solo | 1.001 / 0.999 / 0.997 / 1.014 | 1.09x at 2k, 1.12x at 8k, 1.36x at 32k |
| Nemotron 3.5 Lightning | M5 Max | 22/22, resume 2/2, 16 streams equal solo | 0.984 / 0.999 / 0.969 / 0.999 | 1.05x at 2k, 1.00x at 8k, 1.01x at 32k |
| Qwen3.8-27B | M5 Max | 22/22, resume 2/2, 16 streams equal solo | about 0.97 a round (see below) | 1.00x at 2k, 1.03x at 8k, 1.01x at 16k, 1.03x at 32k |
| Qwen3.8-27B | M3 Ultra | 22/22, resume 2/2, 16 streams equal solo | 1.196 / 1.171 / 1.181 / 1.200 | 0.99x at 2k, 8k and 32k |
| Qwen3.8-27B | DGX Spark | 9/9 | 1.12 / 1.13 / 1.12 / 1.13 | 2.5-3x |
| Qwen3.8 Flash Next | DGX Spark | 9/9 | 1.18 / 1.08 / 1.03 / 1.23 | about 9x |
| GLM-5.3-Flash | two DGX Sparks | 3/3 | 1.088 / 1.012 / 0.973 / 1.011 |
Where 0.3.5 is still short:
- Flash Next prompt processing is 0.80-0.84x of a stock oMLX server on the same M3 Ultra, and Nemotron's is 0.82-0.85x of mlx_lm at 2k-8k on an M5 Max (level at 32k). 0.3.4.1 had the same gaps; closing them is next.
- GLM on two Sparks decodes code greedily 2.7% slower than 0.3.4.1; the automatic drafter choice now re-probes a drafter it left after three rounds, and the rest is the next fix. Nemotron's code-sampled cell reads 1.6% under 0.3.4.1 at the median.
- Resuming at message starts costs cold prefill on prompts with many messages: 7.8% on a 40-message 18.7k session for the 27B, 4.3% for Nemotron, 2.9% for Flash Next. Single-message prompts are level.
- The 27B on M5 decodes about 3% slower than 0.3.4.1 (each round ~1.3 ms dearer, same tokens) in the new engine path; a bounds check that cost another ~8% was removed before release. The rest follows in 0.3.5.1.
- Before M5, Flash Next decodes with 0.3.4.1's per-row kernels, so a single stream matches 0.3.4.1 exactly; faster kernels at the same bits are next.
Contributors
- @chris247474: request reasoning effort and typed tool parameters (#28).
- @quigles1977: the Flash Next prefill path (#29).
- @jasontitus: Metal kernel signature checks (#33) and the 3- and 2-bit lane kernels (#34).
- @taussoe: GLM's sparse-attention scores for any index head count (#43).
- Thanks to @ivanfioravanti for diagnosing the first-decode stall (#30), @chadhurley25075-png for the concurrency benchmark (#37), and everyone who filed issues and independent benchmarks.
GLM-5.3-Flash on Mac (#9, #35, #39), Gemma 4 (#10) and the CUDA prefill and EXL3 work (#40, #42) follow in the next releases.
Install
tensorfold update installs this release. Or:
pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.5TensorFold 0.3.4.1: prompt processing back to full speed
A fix for 0.3.4: prompt processing is back to MLX's speed on every model, and a resumed conversation still matches the same conversation fed fresh, byte for byte.
What was wrong
0.3.4 processed Nemotron's and Flash Next's prompts through their row-exact decode kernels, so that a conversation resumed from a cached prefix got exactly the bits of a fresh one. That made prompt processing several times slower than MLX's own: about 550 tok/s for Nemotron on an M5 Max, where mlx_lm runs about 2,500-3,000, and about 300 for Flash Next on an M3 Ultra, where MLX runs 760.
What changed
- Every model prefills through MLX's own forward, in chunks on a fixed 2,048-token grid from the start of the prompt. Prompt caches are kept only at grid points, and a reply is prefilled again from the last grid point on the next turn. A chunk's bits then never depend on where a conversation was resumed, so resumed conversations still equal fresh ones.
- Flash Next queues at most two layers ahead of the GPU while it prefills, so its peak memory stays bounded at any context length.
- The server no longer holds the first request while it warms saved system blocks at startup; the warm runs in the background.
- Prompt caches saved by 0.3.4 are not reused, because their key now names the prefill mode, so the first prompt after upgrading is prefilled from scratch.
Measured on this tree
Cold prompts, tok/s, through tensorfold serve:
| Model | Machine | Prompt | 0.3.4 | 0.3.4.1 | MLX |
|---|---|---|---|---|---|
| Nemotron 3.5 Lightning | M5 Max | 8k | ~550 | 3,400 | 2,540 (mlx_lm, same session) |
| 32k | ~540 | 2,456 | 2,448 (mlx_lm, same session) | ||
| Qwen3.8-27B | M5 Max | 2k | 542 | 867 | 812 (mlx_lm) |
| 32k | 510 | 560 | 562 (mlx_lm) | ||
| 64k | 441 | 465 | 467 (mlx_lm) | ||
| Qwen3.8 Flash Next | M3 Ultra | 2k | 305 | 934 | |
| 16k | ~290 | 820 | 759 (MLX, at 22.8k) | ||
| 64k | ~280 | 550 | |||
| 196k | ~255 | 315 |
Flash Next's peak memory stays within 20 GB of its weights up to a 196k-token prompt (134 GB peak on the M3 Ultra). Its prompt speed still falls with context length; a sparse prefill kernel that reads only each row's selected keys is next.
Checks on this tree:
- Every drafted reply equals the same request sent with
"draft": false, 9 of 9 on each model. - Resumed multi-turn conversations equal the same conversations fed fresh, including a three-turn Flash Next chat resumed at nine grid points of a 19k-token first message, with thinking on and off.
- Decode speed is unchanged: Nemotron on an M5 Max within noise of 0.3.4 over 15 seeds a cell, and Flash Next on an M3 Ultra at 135.5 / 111.1 / 142.3 / 122.2 against 134.1 / 111.8 / 137.7 / 117.0 (code sampled / chat sampled / code greedy / chat greedy).
- The test suite passes (370 tests).
Install
tensorfold update installs this release. Or:
pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.4.1TensorFold 0.3.4 (work in progress)
Work in progress. 0.3.4 publishes everything built since 0.3.3 as a working snapshot. It's marked as a pre-release, so tensorfold update keeps offering 0.3.3; install it on purpose with the command at the end. Every number below was measured on this tree. Every drafted reply was checked byte for byte against the same request sent with "draft": false.
Every model on the lanes
TensorFold drafts tokens and verifies them in batched "lanes": several drafted rows go through one forward. Before this release, only Qwen3.8-27B ran on the lane engine; Nemotron and Flash Next decoded on a separate serial engine with narrow drafting windows. Now:
- One engine for everything. Nemotron 3.5 Lightning and Qwen3.8 Flash Next run on the lane engine through family rounds, and the serial engine is gone. Every drafted round verifies at least 2 rows, and the server reports
min_rowsand a token-id hash (token_sha) for every reply. - Exact rows on every M chip. Nemotron has row-exact kernels on Macs without the M5's tensor units: windows of 2 to 16 rows reproduce one-row decoding there. So on M1 to M4 it drafts instead of decoding one token at a time. On the M5, every verify matmul (dense and experts) is now TensorFold's own, so exactness no longer depends on the MLX version.
- Resumed prompts equal fresh ones. Nemotron and Flash Next now process prompts through the same row-exact kernels as decoding, so a resumed conversation reproduces the same conversation served fresh. Before, MLX's chunked prefill could make them differ at near-ties.
- Cheaper drafts. Nemotron's MTP head drafts over a 32,768-id vocabulary and Flash Next's over its 79,592-id list, so a draft step reads a quarter of the head. Both lists are built from public text only, and their corpora are recorded next to them (Nemotron's: CPython's standard library plus this package's own source and docs). Flash Next's chained drafts stay on the GPU.
Qwen3.8-27B on M1 to M4: 1.9 to 4x serial
0.3.3 made drafting exact on Macs without tensor units. 0.3.4 makes it fast:
- A new lane decoder for these chips (
kernels/qwen/dense/v1/row_forward.py): fused glue, stacked projections and the lane recurrence. - A new 4-bit matmul on the simdgroup matrix units every Apple GPU has (
kernels/qwen/dense/v1/simd_qmm.py). It dequantizes each 8x8 weight tile once and multiplies up to 16 rows, and a row's bits never depend on how many rows ride with it. On an M3 Ultra a window of 2 to 8 rows costs about 32 ms, against 25 ms for one row. - Serial decoding, windows and prompts all go through it, so it defines the reference the drafted rounds reproduce.
- Prompts resume only from a 2,048-token grid (
TF_ROW_PREFILL=aligned, the default), so a resumed conversation equals a fresh one at MLX's prompt speed. This fixes a 0.3.3 gap: on M1 to M4 a resumed conversation could differ from a fresh one at near-ties (13 of 25 follow-ups in a stress test). - Tool-calling requests on these chips now draft only as many tokens as pay. Before, they always drafted 7.
M3 Ultra, MLX 0.32.0, 64-token replies, thinking off (code sampled / chat sampled / code greedy / chat greedy):
| tok/s | |
|---|---|
| MLX serial | 38.2 / 38.2 / 39.3 / 39.3 |
| 0.3.3 | 63.9 / 47.0 / 64.4 / 51.6 |
| 0.3.4 | 141.3 / 73.9 / 158.4 / 74.2 |
CUDA: lanes on every default path
An audit of every decode path in the CUDA engines found five default configurations that could decode one token a forward. All five are fixed:
- Flash Next always verifies its first MTP draft, so every round is at least 2 rows.
- The 27B refuses to start without its DFlash2 draft model, and Flash Next without its MTP head, each naming the fix (
--no-draftsstill serves the serial reference). Before, both served mostly one token a round. - GLM-5.3-Flash with
--mtp-drafts 0, or a checkpoint without its MTP head, drafts with DFlash2. Before, the first left the loaded draft model idle and the second crashed.
The Spark speeds are unchanged within 1%. On one DGX Spark, the 27B runs 50.0 / 45.5 / 49.3 / 46.0 tok/s and Flash Next 64.6 / 56.9 / 72.5 / 59.2, with 9/9 drafted equal to serial and prefix reuse exact.
Measured on this tree
Same cells: bench_openai, 64-token replies, thinking off.
| Model | Machine | Code sampled | Chat sampled | Code greedy | Chat greedy | Against |
|---|---|---|---|---|---|---|
| Nemotron 3.5 Lightning | M5 Max | 288.0 | 222.9 | 291.8 | 242.6 | mlx_lm's own server 172.6 / 172.8 / 178.6 / 175.9 |
| Nemotron 3.5 Lightning | M3 Ultra | 295.1 | 277.0 | 332.8 | 278.5 | 0.3.3 served it serially there (~220) |
| Qwen3.8 Flash Next | M3 Ultra | 127.4 | 111.1 | 137.7 | 117.8 | 0.3.3: 122.7 / 109.8 / 141.2 / 115.3 |
| Qwen3.8-27B + DFlash2 | M5 Max | 153.2 | 69.8 | 154.2 | 72.8 | unchanged from 0.3.3 |
| Qwen3.8-27B + DFlash2 | M3 Ultra | 141.3 | 73.9 | 158.4 | 74.2 | MLX serial 38.2 / 38.2 / 39.3 / 39.3 |
Every row passed 9 of 9 drafted replies equal to "draft": false, and resumed chats equal to fresh ones.
Known issues
- Some cells are still under twice the standard. That is our floor. The 27B's chat cells on M1 to M4 are at 1.9x, Nemotron's chat cells at 1.3-1.4x mlx_lm, Flash Next at 1.4-1.7x MLX serial on an M3 Ultra, and several CUDA cells.
- Cold long prompts are slower for Nemotron and Flash Next. Exact prompt processing costs speed there: Flash Next about 281 against 759 tok/s on an M3 Ultra, Nemotron about 650 tok/s on an M5 Max. Prompts resumed from a cached prefix are not affected.
- Concurrent requests are served one at a time. Several requests drafting together in one forward is in progress. The first measurement is 4 concurrent 27B requests at 177 tok/s combined on an M5 Max, all 4 exact, but it isn't in this release.
- Draft trees on M1 to M4 are opt-in (
TF_ROW_ATTENTION=1); they're exact but slower than chains today.
Next experiment: MTP heads and drafters
The lanes are cheap now. On an M5 Max the 27B verifies 16 rows in about 36 ms and 128 rows in 145, and 128 lanes can carry about 620 tokens a second. What limits a single request is how many of those lanes hold right guesses, which comes down to the drafter:
- The 27B: chat takes 3.5-3.95 tokens a round, where DFlash2's own top candidates would allow 5.6-6.1.
- Nemotron: on chat, its MTP head's first 1, 2, 3 and 4 chained drafts are all right 74%, 48%, 31% and 16% of the time.
- Flash Next on a Spark: perfect use of today's MTP drafts still leaves the code cells short of twice vLLM.
Tuning trees, draft counts, copies and scoring on today's drafters moves the M5 by no more than ±2%. So the next experiment is the drafters themselves:
- Fine-tune each model's own MTP head on the model's own output, unrolled several drafts deep the way the engine chains them.
- A TensorFold drafter that returns a whole calibrated tree in one small forward, trained on what the engine actually accepts.
Both keep output byte-identical, since drafts only change speed. The goal is every cell at least twice the standard, on every Mac and on DGX Spark.
Install
pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.4TensorFold 0.3.3
All Apple Silicon chips now support lane batching. Qwen3.8-27B verifies its drafted tokens together in one forward on every M1 to M5 GPU, with output byte-identical to serial decoding.
Before 0.3.3, Macs without the M5's tensor units served the 27B at serial speed: the DFlash2 drafter stopped after its first round. Now:
- M1 to M4 GPUs verify drafted windows of 2 to 8 rows through a new row-exact 4-bit matvec (
kernels/qwen/dense/v1/row_qmv.py). A row gets the same bits whatever the window size, so drafted output still equals serial decoding. - At load the engine checks which window widths reproduce one-row decoding on your Mac and times them. Each round then drafts as many tokens as pay at the request's recent acceptance.
- M5 GPUs keep the lane kernels, and their speed is unchanged.
Measured on an M3 Ultra (MLX 0.32.0), 64-token replies, median of seeds 1234-1238, through tools/bench_openai.py:
| Code, sampled | Chat, sampled | Code, greedy | Chat, greedy | |
|---|---|---|---|---|
Serial (--no-drafts) |
38.2 | 38.2 | 39.3 | 39.3 |
| 0.3.3, DFlash2 drafts | 63.9 | 47.0 | 64.4 | 51.6 |
Every drafted reply equaled the same request sent with "draft": false (9 of 9), and resumed conversations gave identical replies. The 27B recipe has the details.
Also fixed: streamed /v1/completions on the Mac server sent reasoning as an object in "text". Text completions now stream plain text, as their non-streamed reply does.
Update from 0.3.2 with tensorfold update, or from an earlier release with:
pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.3TensorFold 0.3.2
TensorFold 0.3.2 adds two things: GLM-5.3-Flash on Mia-AiLab's EXL3 weights, as an experiment, and tensorfold update.
GLM-5.3-Flash on EXL3 (experimental). The CUDA engine now reads Mia-AiLab/GLM-5.3-Flash-EXL3-TR3-4bpw, the checkpoint vLLM serves in Mia-AiLab's DGX Spark recipe, on two Sparks. On those same weights, through the same OpenAI client (decode tok/s, one stream, 64-token replies, median of seeds 1234-1238):
| Code, sampled | Chat, sampled | Code, greedy | Chat, greedy | |
|---|---|---|---|---|
| vLLM, Mia-AiLab's recipe, MTP=3 | 24.5 | 24.3 | 32.2 | 24.7 |
| TensorFold 0.3.2 | 36.4 | 29.7 | 43.8 | 32.9 |
Replies are byte-identical to TensorFold's own serial decoding, as on every other model. The MLX checkpoint (Vontra/GLM-5.3-Flash-MLX-4bit-MTP) is still the faster way to run GLM here, because only the EXL3 checkpoint's routed experts are 4-bit: each Spark reads 10.7 GB a token from it against 5.0. The recipe has the kernels, the checks and what was not checked. The EXL3 format is ExLlamaV3's (MIT).
tensorfold update. tensorfold update installs the newest release. serve also checks GitHub for a newer release once a day in the background, and prints a line if it finds one. It never delays the server. Turn it off with --no-update-check or TENSORFOLD_NO_UPDATE_CHECK=1. From 0.3.1 or earlier, update once by hand:
pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.2TensorFold 0.3.1
TensorFold 0.3.1 says plainly when a checkpoint has no recipe, before anything downloads.
tensorfold info MODELshows how a checkpoint stores its weights (MLX 4-bit, groups of 64,exl3 (4-bit), ...) and which backends read them.serveandpullrefuse a model type with no family, or weights no engine reads yet (EXL3, NVFP4, GPTQ, AWQ), and point to the recipe book and the runbook.- A checkpoint that isn't one TensorFold is tested with still runs when its format matches, with a note that its speed and quality are unmeasured.
Upgrade: pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.1.