TUFF v5.1.0
TUFF 5.1.0 makes long prompts and long-context generation faster, tightens the local server, and fixes the bundled tuff command.
- Faster prompt processing. Full-attention prefill now uses the TensorOps kernel on every Mac that can build it, including M1 and M2, not only the newest GPUs, and for Gemma 4 26B, 12B, E2B and E4B, Qwen 3.6, and Qwen 3.8 Flash Next. Attention is 8-10x faster on an M2; a 14K-token Gemma E4B prompt went from 770 s to 631 s with identical output.
- Faster long-context generation for Gemma 4 12B and Qwen 3.8 Flash Next. Their decode attention now uses the barrier-free kernel the other models already had: at 110K tokens, 40 ms to 12 ms per layer on the 12B and 46 ms to 6 ms on Flash Next.
- Stricter local server. Like the OpenAI API, an unrecognized request field now returns a 400
unknown_parameternaming it instead of being silently ignored, with a suggestion for near misses (max_token) and a pointer from other servers' fields (chat_template_kwargstoenable_thinking). Unsupported OpenAI parameters and JSON response formats returnunsupported_value.reasoning_effortkeeps working for GPT-OSS. - Fixed: the
tuff promptandtuff servecommands bundled inside TUFF.app crashed with "unable to find bundle named TUFF_TUFFEngine" on Macs without a source checkout. - Long prefills set
AGX_RELAX_CDM_CTXSTORE_TIMEOUT=1so macOS is less likely to kill them while the display is active; set it yourself to override. GPU failures now name the layer and phase that failed.
Thanks to drumih/turbo-fieldfare, whose #159, #171 and #182 these changes build on.
Validation: all 1,640 tests passed locally. Every model shape matched a CPU reference, Gemma E4B and 12B produced token-identical output to 5.0.3, and the release archive passed extraction, code-signature, bundled launcher, checksum, and signed update-feed checks, with the packaged app, tuff prompt and tuff serve run from outside the repository.
Download the macOS arm64 ZIP below, or update through TUFF's built-in updater. The app is ad-hoc signed, not notarized.