Skip to content

TUFF v5.2.0

Choose a tag to compare

@rexmhall09 rexmhall09 released this 29 Sep 04:15
· 26 commits to main since this release

TUFF 5.2.0 makes long prompts much faster on models larger than your Mac's memory, and adds fast attention for sliding-window layers and MiniMax.

  • Prefill chunk size now fits the model and the Mac. Prefill reads each layer's experts again for every chunk of the prompt, so on a model whose experts cannot stay in memory, chunk count is SSD traffic. TUFF now uses 2,048-token chunks for a mixture-of-experts model larger than your Mac's memory (1,024 below 16 GB), 512 for one that fits, and 256 for the dense Gemmas. On a 16 GB M2, a 7,000-token Qwen 3.6 prompt went from 190 s to 86 s with identical output. Flash Next, MiniMax M2.7 and GPT-OSS 120B are in the same class. The app, tuff prompt and tuff serve apply this automatically; the app previously used 128 for every model. Memory is sized from the chunk actually chosen, so models that stay at 256 use exactly what they did before.
  • Faster sliding-window prefill attention. The TensorOps attention kernel now handles sliding windows, the FP16 KV ring and image blocks itself, so Gemma 4's sliding layers leave the slow tiled kernel: 8.3x faster per 26B sliding layer. MiniMax M2.7 gets fast prefill attention for the first time (8.6x at 4K keys). End-to-end gains depend on how much of a prompt's time is attention; on a fanless M2 Air they were within run-to-run noise.
  • TUFFCLI gains --prefill-chunk-max to cap --prefill-chunk-tokens auto, and both TUFFCLI and TUFFServer accept chunk sizes up to 2,048.

Validation: all 1,644 tests passed locally. New CPU-reference tests cover sliding windows, ring wrap, bidirectional image blocks and the MiniMax shape. Gemma 4 26B and E4B produced token-identical output to 5.1.0, image prompts and MiniMax ran, and the release archive passed extraction, code-signature, bundled launcher, checksum, and signed update-feed checks, with tuff serve and tuff prompt run from outside the repository.

Download the macOS arm64 ZIP below, or update through TUFF's built-in updater. The app is ad-hoc signed, not notarized.