Skip to content

TensorFold 0.3.5: concurrent requests, resume at message starts, memory that fits

Choose a tag to compare

@ashhart ashhart released this 27 Sep 23:13
· 59 commits to main since this release

Concurrent requests now share each verification round, follow-up turns resume from the start of their newest messages, and the server keeps its whole memory footprint inside what the Mac can hold. Every reply is still byte-identical to the model's own serial decode, and each concurrent stream equals its solo run.

What's new

  • Concurrency. --parallel auto is on by default: requests share one verification round, so aggregate throughput rises with the number of streams. On Metal and CUDA, every stream's reply equals the same request served alone.
  • Resume at message starts. Prompts are chunked where replies begin, so a follow-up turn prefills only its latest reply and new messages, with output identical to a fresh prompt. On a 12-turn agent session with the 27B, first-token time summed to 14.9 s instead of 31.9 s. --prefill-grid is gone.
  • Memory. The whole process stays inside 70% of RAM, an omitted --context defaults to the window the machine can hold (printed at startup), and a prompt past it gets a clear 400 instead of pushing macOS into swap. The README lists tested windows per memory size.
  • Flash Next prefill. @quigles1977's prefill path (#29) with sparse prompt attention: cold prompts 1.1-1.4x faster than 0.3.4.1 on an M3 Ultra, and no pause before the first decoded token after a long prompt.
  • 2- to 8-bit weights. @jasontitus's 3- and 2-bit lane kernels (#34), extended to 5-, 6- and 8-bit, so mixed-precision 27B checkpoints decode fully on the lanes. Checkpoints the lanes can't read are refused from config.json before any download.
  • CUDA. FP8 prefill and shared expert kernels (prefill 2.5-3x faster than 0.3.4.1 for the 27B and about 9x for Flash Next on one DGX Spark), concurrent streams for the 27B and Flash Next, admission from available memory before loading, and the affordable native context by default.
  • API. Raw /v1/completions prompts (no chat template, no think block), ignore_eos and stop strings, request reasoning_effort and typed tool arguments (@chris247474, #28), developer messages as system instructions, parallel_tool_calls: false, tools with null parameters, and cancellation when a client disconnects. An explicit max_tokens that doesn't fit the context now gets a 400; 0.3.4.1 capped it.

Fixed

#19, #20, #21, #23, #25, #26, #27, #31, #32, #36, #46.

Measured on this release

Paired against 0.3.4.1 on the same machine, cold prompts, thinking off; decode cells are 64-token replies (code sampled / chat sampled / code greedy / chat greedy).

Model Machine Exact Decode against 0.3.4.1 Prefill against 0.3.4.1
Qwen3.8 Flash Next M3 Ultra 22/22, resume 2/2, 16 streams equal solo 1.001 / 0.999 / 0.997 / 1.014 1.09x at 2k, 1.12x at 8k, 1.36x at 32k
Nemotron 3.5 Lightning M5 Max 22/22, resume 2/2, 16 streams equal solo 0.984 / 0.999 / 0.969 / 0.999 1.05x at 2k, 1.00x at 8k, 1.01x at 32k
Qwen3.8-27B M5 Max 22/22, resume 2/2, 16 streams equal solo about 0.97 a round (see below) 1.00x at 2k, 1.03x at 8k, 1.01x at 16k, 1.03x at 32k
Qwen3.8-27B M3 Ultra 22/22, resume 2/2, 16 streams equal solo 1.196 / 1.171 / 1.181 / 1.200 0.99x at 2k, 8k and 32k
Qwen3.8-27B DGX Spark 9/9 1.12 / 1.13 / 1.12 / 1.13 2.5-3x
Qwen3.8 Flash Next DGX Spark 9/9 1.18 / 1.08 / 1.03 / 1.23 about 9x
GLM-5.3-Flash two DGX Sparks 3/3 1.088 / 1.012 / 0.973 / 1.011

Where 0.3.5 is still short:

  • Flash Next prompt processing is 0.80-0.84x of a stock oMLX server on the same M3 Ultra, and Nemotron's is 0.82-0.85x of mlx_lm at 2k-8k on an M5 Max (level at 32k). 0.3.4.1 had the same gaps; closing them is next.
  • GLM on two Sparks decodes code greedily 2.7% slower than 0.3.4.1; the automatic drafter choice now re-probes a drafter it left after three rounds, and the rest is the next fix. Nemotron's code-sampled cell reads 1.6% under 0.3.4.1 at the median.
  • Resuming at message starts costs cold prefill on prompts with many messages: 7.8% on a 40-message 18.7k session for the 27B, 4.3% for Nemotron, 2.9% for Flash Next. Single-message prompts are level.
  • The 27B on M5 decodes about 3% slower than 0.3.4.1 (each round ~1.3 ms dearer, same tokens) in the new engine path; a bounds check that cost another ~8% was removed before release. The rest follows in 0.3.5.1.
  • Before M5, Flash Next decodes with 0.3.4.1's per-row kernels, so a single stream matches 0.3.4.1 exactly; faster kernels at the same bits are next.

Contributors

GLM-5.3-Flash on Mac (#9, #35, #39), Gemma 4 (#10) and the CUDA prefill and EXL3 work (#40, #42) follow in the next releases.

Install

tensorfold update installs this release. Or:

pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.5