TensorFold 0.3.3
All Apple Silicon chips now support lane batching. Qwen3.8-27B verifies its drafted tokens together in one forward on every M1 to M5 GPU, with output byte-identical to serial decoding.
Before 0.3.3, Macs without the M5's tensor units served the 27B at serial speed: the DFlash2 drafter stopped after its first round. Now:
- M1 to M4 GPUs verify drafted windows of 2 to 8 rows through a new row-exact 4-bit matvec (
kernels/qwen/dense/v1/row_qmv.py). A row gets the same bits whatever the window size, so drafted output still equals serial decoding. - At load the engine checks which window widths reproduce one-row decoding on your Mac and times them. Each round then drafts as many tokens as pay at the request's recent acceptance.
- M5 GPUs keep the lane kernels, and their speed is unchanged.
Measured on an M3 Ultra (MLX 0.32.0), 64-token replies, median of seeds 1234-1238, through tools/bench_openai.py:
| Code, sampled | Chat, sampled | Code, greedy | Chat, greedy | |
|---|---|---|---|---|
Serial (--no-drafts) |
38.2 | 38.2 | 39.3 | 39.3 |
| 0.3.3, DFlash2 drafts | 63.9 | 47.0 | 64.4 | 51.6 |
Every drafted reply equaled the same request sent with "draft": false (9 of 9), and resumed conversations gave identical replies. The 27B recipe has the details.
Also fixed: streamed /v1/completions on the Mac server sent reasoning as an object in "text". Text completions now stream plain text, as their non-streamed reply does.
Update from 0.3.2 with tensorfold update, or from an earlier release with:
pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.3