TUFF v6.0.0
TUFF 6.0.0 rebuilds prefill around batched tensor-core matrix multiplies and reads mixture-of-experts weights ahead in decode. Long prompts on dense models are an order of magnitude faster, and every streaming MoE model decodes faster.
- Batched prefill. Within each prompt chunk, the dense and shared MLPs, every routed expert, Qwen 3.8 Flash Next's group-32 projections, and GPT-OSS's MXFP4 experts and BF16 attention projections now run as MPP matrix multiplies instead of one matrix-vector product per token. Each routed expert gets a tile sized to its rows (8, 16, 32 or 64), and every tile dequantizes 8 weights per load. GPT-OSS applies RoPE and picks experts once per chunk. Blocks under 32 tokens, including speculative verification, keep the decode kernels.
- Next-layer expert prefetch in decode. Each layer's input also goes through the next layer's router, and the experts it predicts are read from SSD while the GPU works. It turns itself off if fewer than 35% of guesses are used; a guess only decides what is read early, so output is unchanged.
- Faster GPT-OSS MXFP4 decode, about four times faster per byte.
Prefill on a 16 GB M2, alternating the packaged 5.2.0 and 6.0.0 apps on the same day:
| Model | Prompt | 5.2.0 | 6.0.0 |
|---|---|---|---|
| Gemma 4 E4B | 4,156 tokens | 145 s | 19 s |
| Gemma 4 12B QAT | 1,004 tokens | 360 s | 24 s |
| GPT-OSS 20B | 973 tokens | 127 s | 34 s |
| Qwen 3.6 35B-A3B | 4,087 tokens | 169 s | 155 s |
| Gemma 4 26B-A4B | 4,160 tokens | 164 s | 127 s |
| Qwen 3.8 Flash Next | 983 tokens | 156 s | 118 s |
The last three stream experts from SSD for every chunk on a 16 GB Mac, so their prompt time is mostly reads.
64-token decode, best of two alternating runs, identical output:
| Model | 5.2.0 | 6.0.0 | Guesses used |
|---|---|---|---|
| Qwen 3.6 35B-A3B | 7.3 tok/s | 11.4 tok/s | 84% |
| Gemma 4 26B-A4B | 8.5 tok/s | 10.4 tok/s | 71% |
| Qwen 3.8 Flash Next | 2.4 tok/s | 2.6 tok/s | 69% |
| GPT-OSS 20B | 4.9 tok/s | 6.4 tok/s | 90% |
Output matches 5.2.0 on these prompts except on Qwen 3.6 and Gemma 4 E2B, where the batched kernels' summation order changes a word or two of a long-prompt continuation.
Validation: all 1,651 tests passed locally, including new CPU-reference tests for every tile size, unaligned weights, and the GPT-OSS batched paths. All 9 catalog models answered the release check from the packaged app, and the release archive passed extraction, code-signature, bundled launcher, checksum and update-feed checks, with tuff serve and tuff prompt run from outside the repository.
Download the macOS arm64 ZIP below, or update through TUFF's built-in updater. The app is ad-hoc signed, not notarized.