TensorFold 0.3.5: concurrent requests, resume at message starts, memory that fits
Concurrent requests now share each verification round, follow-up turns resume from the start of their newest messages, and the server keeps its whole memory footprint inside what the Mac can hold. Every reply is still byte-identical to the model's own serial decode, and each concurrent stream equals its solo run.
What's new
- Concurrency.
--parallel autois on by default: requests share one verification round, so aggregate throughput rises with the number of streams. On Metal and CUDA, every stream's reply equals the same request served alone. - Resume at message starts. Prompts are chunked where replies begin, so a follow-up turn prefills only its latest reply and new messages, with output identical to a fresh prompt. On a 12-turn agent session with the 27B, first-token time summed to 14.9 s instead of 31.9 s.
--prefill-gridis gone. - Memory. The whole process stays inside 70% of RAM, an omitted
--contextdefaults to the window the machine can hold (printed at startup), and a prompt past it gets a clear 400 instead of pushing macOS into swap. The README lists tested windows per memory size. - Flash Next prefill. @quigles1977's prefill path (#29) with sparse prompt attention: cold prompts 1.1-1.4x faster than 0.3.4.1 on an M3 Ultra, and no pause before the first decoded token after a long prompt.
- 2- to 8-bit weights. @jasontitus's 3- and 2-bit lane kernels (#34), extended to 5-, 6- and 8-bit, so mixed-precision 27B checkpoints decode fully on the lanes. Checkpoints the lanes can't read are refused from
config.jsonbefore any download. - CUDA. FP8 prefill and shared expert kernels (prefill 2.5-3x faster than 0.3.4.1 for the 27B and about 9x for Flash Next on one DGX Spark), concurrent streams for the 27B and Flash Next, admission from available memory before loading, and the affordable native context by default.
- API. Raw
/v1/completionsprompts (no chat template, no think block),ignore_eosandstopstrings, requestreasoning_effortand typed tool arguments (@chris247474, #28),developermessages as system instructions,parallel_tool_calls: false, tools with null parameters, and cancellation when a client disconnects. An explicitmax_tokensthat doesn't fit the context now gets a 400; 0.3.4.1 capped it.
Fixed
#19, #20, #21, #23, #25, #26, #27, #31, #32, #36, #46.
Measured on this release
Paired against 0.3.4.1 on the same machine, cold prompts, thinking off; decode cells are 64-token replies (code sampled / chat sampled / code greedy / chat greedy).
| Model | Machine | Exact | Decode against 0.3.4.1 | Prefill against 0.3.4.1 |
|---|---|---|---|---|
| Qwen3.8 Flash Next | M3 Ultra | 22/22, resume 2/2, 16 streams equal solo | 1.001 / 0.999 / 0.997 / 1.014 | 1.09x at 2k, 1.12x at 8k, 1.36x at 32k |
| Nemotron 3.5 Lightning | M5 Max | 22/22, resume 2/2, 16 streams equal solo | 0.984 / 0.999 / 0.969 / 0.999 | 1.05x at 2k, 1.00x at 8k, 1.01x at 32k |
| Qwen3.8-27B | M5 Max | 22/22, resume 2/2, 16 streams equal solo | about 0.97 a round (see below) | 1.00x at 2k, 1.03x at 8k, 1.01x at 16k, 1.03x at 32k |
| Qwen3.8-27B | M3 Ultra | 22/22, resume 2/2, 16 streams equal solo | 1.196 / 1.171 / 1.181 / 1.200 | 0.99x at 2k, 8k and 32k |
| Qwen3.8-27B | DGX Spark | 9/9 | 1.12 / 1.13 / 1.12 / 1.13 | 2.5-3x |
| Qwen3.8 Flash Next | DGX Spark | 9/9 | 1.18 / 1.08 / 1.03 / 1.23 | about 9x |
| GLM-5.3-Flash | two DGX Sparks | 3/3 | 1.088 / 1.012 / 0.973 / 1.011 |
Where 0.3.5 is still short:
- Flash Next prompt processing is 0.80-0.84x of a stock oMLX server on the same M3 Ultra, and Nemotron's is 0.82-0.85x of mlx_lm at 2k-8k on an M5 Max (level at 32k). 0.3.4.1 had the same gaps; closing them is next.
- GLM on two Sparks decodes code greedily 2.7% slower than 0.3.4.1; the automatic drafter choice now re-probes a drafter it left after three rounds, and the rest is the next fix. Nemotron's code-sampled cell reads 1.6% under 0.3.4.1 at the median.
- Resuming at message starts costs cold prefill on prompts with many messages: 7.8% on a 40-message 18.7k session for the 27B, 4.3% for Nemotron, 2.9% for Flash Next. Single-message prompts are level.
- The 27B on M5 decodes about 3% slower than 0.3.4.1 (each round ~1.3 ms dearer, same tokens) in the new engine path; a bounds check that cost another ~8% was removed before release. The rest follows in 0.3.5.1.
- Before M5, Flash Next decodes with 0.3.4.1's per-row kernels, so a single stream matches 0.3.4.1 exactly; faster kernels at the same bits are next.
Contributors
- @chris247474: request reasoning effort and typed tool parameters (#28).
- @quigles1977: the Flash Next prefill path (#29).
- @jasontitus: Metal kernel signature checks (#33) and the 3- and 2-bit lane kernels (#34).
- @taussoe: GLM's sparse-attention scores for any index head count (#43).
- Thanks to @ivanfioravanti for diagnosing the first-decode stall (#30), @chadhurley25075-png for the concurrency benchmark (#37), and everyone who filed issues and independent benchmarks.
GLM-5.3-Flash on Mac (#9, #35, #39), Gemma 4 (#10) and the CUDA prefill and EXL3 work (#40, #42) follow in the next releases.
Install
tensorfold update installs this release. Or:
pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.5