TensorFold 0.3.4 (work in progress)
Pre-releaseWork in progress. 0.3.4 publishes everything built since 0.3.3 as a working snapshot. It's marked as a pre-release, so tensorfold update keeps offering 0.3.3; install it on purpose with the command at the end. Every number below was measured on this tree. Every drafted reply was checked byte for byte against the same request sent with "draft": false.
Every model on the lanes
TensorFold drafts tokens and verifies them in batched "lanes": several drafted rows go through one forward. Before this release, only Qwen3.8-27B ran on the lane engine; Nemotron and Flash Next decoded on a separate serial engine with narrow drafting windows. Now:
- One engine for everything. Nemotron 3.5 Lightning and Qwen3.8 Flash Next run on the lane engine through family rounds, and the serial engine is gone. Every drafted round verifies at least 2 rows, and the server reports
min_rowsand a token-id hash (token_sha) for every reply. - Exact rows on every M chip. Nemotron has row-exact kernels on Macs without the M5's tensor units: windows of 2 to 16 rows reproduce one-row decoding there. So on M1 to M4 it drafts instead of decoding one token at a time. On the M5, every verify matmul (dense and experts) is now TensorFold's own, so exactness no longer depends on the MLX version.
- Resumed prompts equal fresh ones. Nemotron and Flash Next now process prompts through the same row-exact kernels as decoding, so a resumed conversation reproduces the same conversation served fresh. Before, MLX's chunked prefill could make them differ at near-ties.
- Cheaper drafts. Nemotron's MTP head drafts over a 32,768-id vocabulary and Flash Next's over its 79,592-id list, so a draft step reads a quarter of the head. Both lists are built from public text only, and their corpora are recorded next to them (Nemotron's: CPython's standard library plus this package's own source and docs). Flash Next's chained drafts stay on the GPU.
Qwen3.8-27B on M1 to M4: 1.9 to 4x serial
0.3.3 made drafting exact on Macs without tensor units. 0.3.4 makes it fast:
- A new lane decoder for these chips (
kernels/qwen/dense/v1/row_forward.py): fused glue, stacked projections and the lane recurrence. - A new 4-bit matmul on the simdgroup matrix units every Apple GPU has (
kernels/qwen/dense/v1/simd_qmm.py). It dequantizes each 8x8 weight tile once and multiplies up to 16 rows, and a row's bits never depend on how many rows ride with it. On an M3 Ultra a window of 2 to 8 rows costs about 32 ms, against 25 ms for one row. - Serial decoding, windows and prompts all go through it, so it defines the reference the drafted rounds reproduce.
- Prompts resume only from a 2,048-token grid (
TF_ROW_PREFILL=aligned, the default), so a resumed conversation equals a fresh one at MLX's prompt speed. This fixes a 0.3.3 gap: on M1 to M4 a resumed conversation could differ from a fresh one at near-ties (13 of 25 follow-ups in a stress test). - Tool-calling requests on these chips now draft only as many tokens as pay. Before, they always drafted 7.
M3 Ultra, MLX 0.32.0, 64-token replies, thinking off (code sampled / chat sampled / code greedy / chat greedy):
| tok/s | |
|---|---|
| MLX serial | 38.2 / 38.2 / 39.3 / 39.3 |
| 0.3.3 | 63.9 / 47.0 / 64.4 / 51.6 |
| 0.3.4 | 141.3 / 73.9 / 158.4 / 74.2 |
CUDA: lanes on every default path
An audit of every decode path in the CUDA engines found five default configurations that could decode one token a forward. All five are fixed:
- Flash Next always verifies its first MTP draft, so every round is at least 2 rows.
- The 27B refuses to start without its DFlash2 draft model, and Flash Next without its MTP head, each naming the fix (
--no-draftsstill serves the serial reference). Before, both served mostly one token a round. - GLM-5.3-Flash with
--mtp-drafts 0, or a checkpoint without its MTP head, drafts with DFlash2. Before, the first left the loaded draft model idle and the second crashed.
The Spark speeds are unchanged within 1%. On one DGX Spark, the 27B runs 50.0 / 45.5 / 49.3 / 46.0 tok/s and Flash Next 64.6 / 56.9 / 72.5 / 59.2, with 9/9 drafted equal to serial and prefix reuse exact.
Measured on this tree
Same cells: bench_openai, 64-token replies, thinking off.
| Model | Machine | Code sampled | Chat sampled | Code greedy | Chat greedy | Against |
|---|---|---|---|---|---|---|
| Nemotron 3.5 Lightning | M5 Max | 288.0 | 222.9 | 291.8 | 242.6 | mlx_lm's own server 172.6 / 172.8 / 178.6 / 175.9 |
| Nemotron 3.5 Lightning | M3 Ultra | 295.1 | 277.0 | 332.8 | 278.5 | 0.3.3 served it serially there (~220) |
| Qwen3.8 Flash Next | M3 Ultra | 127.4 | 111.1 | 137.7 | 117.8 | 0.3.3: 122.7 / 109.8 / 141.2 / 115.3 |
| Qwen3.8-27B + DFlash2 | M5 Max | 153.2 | 69.8 | 154.2 | 72.8 | unchanged from 0.3.3 |
| Qwen3.8-27B + DFlash2 | M3 Ultra | 141.3 | 73.9 | 158.4 | 74.2 | MLX serial 38.2 / 38.2 / 39.3 / 39.3 |
Every row passed 9 of 9 drafted replies equal to "draft": false, and resumed chats equal to fresh ones.
Known issues
- Some cells are still under twice the standard. That is our floor. The 27B's chat cells on M1 to M4 are at 1.9x, Nemotron's chat cells at 1.3-1.4x mlx_lm, Flash Next at 1.4-1.7x MLX serial on an M3 Ultra, and several CUDA cells.
- Cold long prompts are slower for Nemotron and Flash Next. Exact prompt processing costs speed there: Flash Next about 281 against 759 tok/s on an M3 Ultra, Nemotron about 650 tok/s on an M5 Max. Prompts resumed from a cached prefix are not affected.
- Concurrent requests are served one at a time. Several requests drafting together in one forward is in progress. The first measurement is 4 concurrent 27B requests at 177 tok/s combined on an M5 Max, all 4 exact, but it isn't in this release.
- Draft trees on M1 to M4 are opt-in (
TF_ROW_ATTENTION=1); they're exact but slower than chains today.
Next experiment: MTP heads and drafters
The lanes are cheap now. On an M5 Max the 27B verifies 16 rows in about 36 ms and 128 rows in 145, and 128 lanes can carry about 620 tokens a second. What limits a single request is how many of those lanes hold right guesses, which comes down to the drafter:
- The 27B: chat takes 3.5-3.95 tokens a round, where DFlash2's own top candidates would allow 5.6-6.1.
- Nemotron: on chat, its MTP head's first 1, 2, 3 and 4 chained drafts are all right 74%, 48%, 31% and 16% of the time.
- Flash Next on a Spark: perfect use of today's MTP drafts still leaves the code cells short of twice vLLM.
Tuning trees, draft counts, copies and scoring on today's drafters moves the M5 by no more than ±2%. So the next experiment is the drafters themselves:
- Fine-tune each model's own MTP head on the model's own output, unrolled several drafts deep the way the engine chains them.
- A TensorFold drafter that returns a whole calibrated tree in one small forward, trained on what the engine actually accepts.
Both keep output byte-identical, since drafts only change speed. The goal is every cell at least twice the standard, on every Mac and on DGX Spark.
Install
pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.4