Skip to content

TensorFold 0.3.3

Choose a tag to compare

@ashhart ashhart released this 26 Sep 21:11
· 71 commits to main since this release
b08e014

All Apple Silicon chips now support lane batching. Qwen3.8-27B verifies its drafted tokens together in one forward on every M1 to M5 GPU, with output byte-identical to serial decoding.

Before 0.3.3, Macs without the M5's tensor units served the 27B at serial speed: the DFlash2 drafter stopped after its first round. Now:

  • M1 to M4 GPUs verify drafted windows of 2 to 8 rows through a new row-exact 4-bit matvec (kernels/qwen/dense/v1/row_qmv.py). A row gets the same bits whatever the window size, so drafted output still equals serial decoding.
  • At load the engine checks which window widths reproduce one-row decoding on your Mac and times them. Each round then drafts as many tokens as pay at the request's recent acceptance.
  • M5 GPUs keep the lane kernels, and their speed is unchanged.

Measured on an M3 Ultra (MLX 0.32.0), 64-token replies, median of seeds 1234-1238, through tools/bench_openai.py:

Code, sampled Chat, sampled Code, greedy Chat, greedy
Serial (--no-drafts) 38.2 38.2 39.3 39.3
0.3.3, DFlash2 drafts 63.9 47.0 64.4 51.6

Every drafted reply equaled the same request sent with "draft": false (9 of 9), and resumed conversations gave identical replies. The 27B recipe has the details.

Also fixed: streamed /v1/completions on the Mac server sent reasoning as an object in "text". Text completions now stream plain text, as their non-streamed reply does.

Update from 0.3.2 with tensorfold update, or from an earlier release with:

pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.3