Skip to content

TUFF v6.0.0

Choose a tag to compare

@rexmhall09 rexmhall09 released this 30 Sep 02:10
· 15 commits to main since this release

TUFF 6.0.0 rebuilds prefill around batched tensor-core matrix multiplies and reads mixture-of-experts weights ahead in decode. Long prompts on dense models are an order of magnitude faster, and every streaming MoE model decodes faster.

  • Batched prefill. Within each prompt chunk, the dense and shared MLPs, every routed expert, Qwen 3.8 Flash Next's group-32 projections, and GPT-OSS's MXFP4 experts and BF16 attention projections now run as MPP matrix multiplies instead of one matrix-vector product per token. Each routed expert gets a tile sized to its rows (8, 16, 32 or 64), and every tile dequantizes 8 weights per load. GPT-OSS applies RoPE and picks experts once per chunk. Blocks under 32 tokens, including speculative verification, keep the decode kernels.
  • Next-layer expert prefetch in decode. Each layer's input also goes through the next layer's router, and the experts it predicts are read from SSD while the GPU works. It turns itself off if fewer than 35% of guesses are used; a guess only decides what is read early, so output is unchanged.
  • Faster GPT-OSS MXFP4 decode, about four times faster per byte.

Prefill on a 16 GB M2, alternating the packaged 5.2.0 and 6.0.0 apps on the same day:

Model Prompt 5.2.0 6.0.0
Gemma 4 E4B 4,156 tokens 145 s 19 s
Gemma 4 12B QAT 1,004 tokens 360 s 24 s
GPT-OSS 20B 973 tokens 127 s 34 s
Qwen 3.6 35B-A3B 4,087 tokens 169 s 155 s
Gemma 4 26B-A4B 4,160 tokens 164 s 127 s
Qwen 3.8 Flash Next 983 tokens 156 s 118 s

The last three stream experts from SSD for every chunk on a 16 GB Mac, so their prompt time is mostly reads.

64-token decode, best of two alternating runs, identical output:

Model 5.2.0 6.0.0 Guesses used
Qwen 3.6 35B-A3B 7.3 tok/s 11.4 tok/s 84%
Gemma 4 26B-A4B 8.5 tok/s 10.4 tok/s 71%
Qwen 3.8 Flash Next 2.4 tok/s 2.6 tok/s 69%
GPT-OSS 20B 4.9 tok/s 6.4 tok/s 90%

Output matches 5.2.0 on these prompts except on Qwen 3.6 and Gemma 4 E2B, where the batched kernels' summation order changes a word or two of a long-prompt continuation.

Validation: all 1,651 tests passed locally, including new CPU-reference tests for every tile size, unaligned weights, and the GPT-OSS batched paths. All 9 catalog models answered the release check from the packaged app, and the release archive passed extraction, code-signature, bundled launcher, checksum and update-feed checks, with tuff serve and tuff prompt run from outside the repository.

Download the macOS arm64 ZIP below, or update through TUFF's built-in updater. The app is ad-hoc signed, not notarized.