Releases: rexmhall09/TUFF
Release list
TUFF 7.1.0
TUFF 7.1.0 keeps one local server, the model-routing Background API introduced in 7.0.0, and makes it work with coding agents such as oh-my-pi.
- One server. The Background API and
tuff serverun the same server: every installed model on one loopback endpoint, loaded when a request names it and unloaded when idle.defaultmeans your chosen model. Each model runs with its catalog context, expert-cache and prefill settings for your Mac. - Server screen. Rebuilt around the Background API: an on/off switch, the endpoint, default model, unload delay and port, plus the loaded model with its unload countdown, active and queued requests, and Open Log. Clone builds show the command to run the same server in Terminal.
- Qwen with oh-my-pi. Qwen clients send
preserve_thinkingandchat_template_kwargs, which the server used to refuse with a 400. It now acceptschat_template_kwargsholdingenable_thinkingorpreserve_thinkingand still refuses any other template argument.
Removed: the app's Start/Stop local server, which served only Chat's loaded model, and fixed-model serving with tuff serve --model / TUFFServer --model and its tuning flags. A removed flag prints what to use instead. --all-models is still accepted and has no effect.
Known limits: oh-my-pi's full default toolset fails on Gemma models because its task tool uses a schema Gemma's tool template cannot represent; after a Qwen tool call the next request re-reads the whole prompt; Qwen3.8 Flash Next's 2,048-token context only just fits oh-my-pi's prompt without tools.
Inference kernels and model packs are unchanged from 7.0.0.
Validation on one 16 GB M2 MacBook Air: 1,698 Swift tests and the model-free gate, packaging and isolated updater fixtures; the packaged CLI; and oh-my-pi through the packaged server on Gemma 4 E4B and 26B-A4B, Qwen3.6 and Qwen3.8 Flash Next, including tool calls. Other Macs were not tested. See the release validation for details and limits.
The app is arm64 and ad-hoc signed, not notarized. The release includes the ZIP, a SHA-256 checksum and a signed Sparkle update feed.
AI assistance: Claude implemented, tested and documented these changes, and OpenAI Codex reviewed them.
TUFF 7.0.0
TUFF 7.0.0 adds a Background API that keeps a local OpenAI-compatible endpoint running while the app is closed, plus a release recovery system and better tools for reporting bugs and contributing.
- Background API. Turn it on in the Server screen and allow the login item. Any installed model can be requested by name, or
defaultfor the model you choose. A model loads when a request arrives and unloads after an idle delay you pick, down to immediately after each response. One request runs at a time; requests for the loaded model go first, with a limit so others are not starved. Nothing unloads while a request is running or queued. Chat and the API share one memory budget, so the API declines a second model that would not fit beside Chat's. Clone builds can run the same router withtuff serve --all-models. - Recovery. Updates now require a signed feed and are verified before extraction. A defective release can be withdrawn without deleting tags or assets, and a newer recovery release can restore known-good code. The app checks that a recovery can read your chats and settings before installing it, and TUFF > Check for Recovery Update reports what is published. See recovery help, including steps for an app that will not open.
- Bug reports. Help > Report a Bug opens the bug form with your version and Mac filled in. An optional diagnostic summary is previewed first and never includes chats, images, credentials or file paths.
- Contributing. New issue forms and labels,
Scripts/check.shfor every model-free check, CI on PRs and pushes to main, and a repository Codex review skill.
Inference kernels and model packs are unchanged from 6.1.0. Kernel experiments are written up in the inference notes; none qualified, and no speed change is claimed.
Validation on one 16 GB M2 MacBook Air: 1,716 Swift tests and 45 harness tests, packaging and isolated updater fixtures; all nine models and six image companions on the packaged CLI; 18 app-service and 18 HTTP-server checks; the router and login item through the official openai Python client; and paired 6.1.0/7.0.0 comparisons. Other Macs were not tested. See the release validation for details and limits.
The app is arm64 and ad-hoc signed, not notarized. The release includes the ZIP, a SHA-256 checksum and a signed Sparkle update feed.
AI assistance: OpenAI Codex and Claude implemented, reviewed, tested and documented these changes.
TUFF 6.1.0
TUFF 6.1.0 extends the efficient GPU sampler to Flash Next top-k 20 and MiniMax top-k 40, and adds repeatable inference comparisons and clearer expert-cache diagnostics.
- Separate demand and prediction histories. Report first demand hits on prefetched records, unused evictions and exposed waits while retaining the validated eviction ranking.
- Retain the synchronized expert streamer for background lookahead, with a startup-only off switch for comparisons.
- Account for the extra cache metadata in shared app, CLI and server memory budgets. Keep qualified cache, grouping and chunk defaults.
- Record short/long, greedy/default sampled and cold/warm comparisons, resolved settings, model identities and available machine-state observations.
Isolated sampler wall medians improved from 23.20 to 1.27 ms for Flash Next and 58.87 to 1.03 ms for MiniMax. Full-inference results were mixed. Demand-only eviction, unused-prefetch priority and grouping experiments were removed; aging was not qualified end to end. No general model speedup is claimed.
Validation: 1,682 canonical tests across 12 targets, 29 harness/reporting regressions, repository checks and release build/archive checks. All nine text models and six image companions passed packaged CLI checks, with nine app-service and 18 HTTP-server passes. The warm comparison passed 96 requests with matching visible output in all 48 pairs. Four sleep-interrupted timing observations remain labeled, with awake replacements. All model work ran serially on one 16 GB M2 MacBook Air; other hardware was not tested.
The app is arm64 and ad-hoc signed, not notarized. The release includes the ZIP, a SHA-256 checksum and a signed Sparkle update feed. See the release validation and model report for settings, all repetitions and limitations.
AI assistance: OpenAI Codex helped implement, review, test and document these changes. Results came from the checks described above.
TUFF v6.0.2
TUFF 6.0.2 corrects inference memory planning and performance reporting.
- Count MoE expert-cache slots across every layer, including GPT-OSS. Cross-check all six MoE layouts against their manifests and actual allocations.
- Include chunk-dependent prefill scratch, sliding-window rings and conservative growth reserves in the shared memory plan used by the app, CLI and server.
- Report demand and prefetched expert records, logical bytes and failures separately. Measure exposed prefetch waits and remove the unsupported GPU-wait residual label.
- Keep every benchmark repetition with median and spread, resolved settings and available machine-state observations. Restore all measured README benchmarks and update the emoji comparison with sourced capabilities and current TUFF features.
Validation: 1,662 canonical tests across 12 targets, 11 benchmark-reporting regression tests, repository checks, release build and archive checks. The packaged CLI passed 27 text attempts across all nine supported models and all six image companions. Flash Next and Gemma 26B also passed packaged app decode-service and HTTP-server checks. All real-model work ran serially on a 16 GB M2 MacBook Air.
These are correctness smoke checks, not general model-quality or sustained-performance qualification. Other chips and memory capacities were not tested. Logical reads include OS-cache hits and do not measure physical SSD traffic; overlapping phase timings are not additive wall-clock totals. No inference speedup is claimed.
The app is arm64 and ad-hoc signed, not notarized. The archive has a SHA-256 checksum and a signed Sparkle update feed. See the repository's model and release validation reports for all repetitions, settings and limitations.
AI assistance: Codex helped implement, review and document these changes. Validation results were collected from the local checks described above.
TUFF v6.0.1
TUFF 6.0.1 fixes Gemma 4 E2B with thinking off.
- Gemma 4 E2B answers directly when thinking is off. Its chat template, like E4B's, opens the model turn and adds nothing when thinking is off. TUFF closed an empty thought channel there instead, which E2B answered by narrating its reasoning as the reply. Full prompts, text follow-ups and image follow-ups now all use each Gemma variant's own template. Thinking on is unchanged.
- The site and README screenshot are retaken in the app's default accent.
Validation: all 1,653 tests passed locally, including new template tests for every Gemma 4 variant. The release archive passed extraction, code-signature, bundled launcher and checksum checks, with tuff prompt run from outside the repository.
Download the macOS arm64 ZIP below, or update through TUFF's built-in updater. The app is ad-hoc signed, not notarized.
TUFF v6.0.0
TUFF 6.0.0 rebuilds prefill around batched tensor-core matrix multiplies and reads mixture-of-experts weights ahead in decode. Long prompts on dense models are an order of magnitude faster, and every streaming MoE model decodes faster.
- Batched prefill. Within each prompt chunk, the dense and shared MLPs, every routed expert, Qwen 3.8 Flash Next's group-32 projections, and GPT-OSS's MXFP4 experts and BF16 attention projections now run as MPP matrix multiplies instead of one matrix-vector product per token. Each routed expert gets a tile sized to its rows (8, 16, 32 or 64), and every tile dequantizes 8 weights per load. GPT-OSS applies RoPE and picks experts once per chunk. Blocks under 32 tokens, including speculative verification, keep the decode kernels.
- Next-layer expert prefetch in decode. Each layer's input also goes through the next layer's router, and the experts it predicts are read from SSD while the GPU works. It turns itself off if fewer than 35% of guesses are used; a guess only decides what is read early, so output is unchanged.
- Faster GPT-OSS MXFP4 decode, about four times faster per byte.
Prefill on a 16 GB M2, alternating the packaged 5.2.0 and 6.0.0 apps on the same day:
| Model | Prompt | 5.2.0 | 6.0.0 |
|---|---|---|---|
| Gemma 4 E4B | 4,156 tokens | 145 s | 19 s |
| Gemma 4 12B QAT | 1,004 tokens | 360 s | 24 s |
| GPT-OSS 20B | 973 tokens | 127 s | 34 s |
| Qwen 3.6 35B-A3B | 4,087 tokens | 169 s | 155 s |
| Gemma 4 26B-A4B | 4,160 tokens | 164 s | 127 s |
| Qwen 3.8 Flash Next | 983 tokens | 156 s | 118 s |
The last three stream experts from SSD for every chunk on a 16 GB Mac, so their prompt time is mostly reads.
64-token decode, best of two alternating runs, identical output:
| Model | 5.2.0 | 6.0.0 | Guesses used |
|---|---|---|---|
| Qwen 3.6 35B-A3B | 7.3 tok/s | 11.4 tok/s | 84% |
| Gemma 4 26B-A4B | 8.5 tok/s | 10.4 tok/s | 71% |
| Qwen 3.8 Flash Next | 2.4 tok/s | 2.6 tok/s | 69% |
| GPT-OSS 20B | 4.9 tok/s | 6.4 tok/s | 90% |
Output matches 5.2.0 on these prompts except on Qwen 3.6 and Gemma 4 E2B, where the batched kernels' summation order changes a word or two of a long-prompt continuation.
Validation: all 1,651 tests passed locally, including new CPU-reference tests for every tile size, unaligned weights, and the GPT-OSS batched paths. All 9 catalog models answered the release check from the packaged app, and the release archive passed extraction, code-signature, bundled launcher, checksum and update-feed checks, with tuff serve and tuff prompt run from outside the repository.
Download the macOS arm64 ZIP below, or update through TUFF's built-in updater. The app is ad-hoc signed, not notarized.
TUFF v5.2.0
TUFF 5.2.0 makes long prompts much faster on models larger than your Mac's memory, and adds fast attention for sliding-window layers and MiniMax.
- Prefill chunk size now fits the model and the Mac. Prefill reads each layer's experts again for every chunk of the prompt, so on a model whose experts cannot stay in memory, chunk count is SSD traffic. TUFF now uses 2,048-token chunks for a mixture-of-experts model larger than your Mac's memory (1,024 below 16 GB), 512 for one that fits, and 256 for the dense Gemmas. On a 16 GB M2, a 7,000-token Qwen 3.6 prompt went from 190 s to 86 s with identical output. Flash Next, MiniMax M2.7 and GPT-OSS 120B are in the same class. The app,
tuff promptandtuff serveapply this automatically; the app previously used 128 for every model. Memory is sized from the chunk actually chosen, so models that stay at 256 use exactly what they did before. - Faster sliding-window prefill attention. The TensorOps attention kernel now handles sliding windows, the FP16 KV ring and image blocks itself, so Gemma 4's sliding layers leave the slow tiled kernel: 8.3x faster per 26B sliding layer. MiniMax M2.7 gets fast prefill attention for the first time (8.6x at 4K keys). End-to-end gains depend on how much of a prompt's time is attention; on a fanless M2 Air they were within run-to-run noise.
TUFFCLIgains--prefill-chunk-maxto cap--prefill-chunk-tokens auto, and bothTUFFCLIandTUFFServeraccept chunk sizes up to 2,048.
Validation: all 1,644 tests passed locally. New CPU-reference tests cover sliding windows, ring wrap, bidirectional image blocks and the MiniMax shape. Gemma 4 26B and E4B produced token-identical output to 5.1.0, image prompts and MiniMax ran, and the release archive passed extraction, code-signature, bundled launcher, checksum, and signed update-feed checks, with tuff serve and tuff prompt run from outside the repository.
Download the macOS arm64 ZIP below, or update through TUFF's built-in updater. The app is ad-hoc signed, not notarized.
TUFF v5.1.0
TUFF 5.1.0 makes long prompts and long-context generation faster, tightens the local server, and fixes the bundled tuff command.
- Faster prompt processing. Full-attention prefill now uses the TensorOps kernel on every Mac that can build it, including M1 and M2, not only the newest GPUs, and for Gemma 4 26B, 12B, E2B and E4B, Qwen 3.6, and Qwen 3.8 Flash Next. Attention is 8-10x faster on an M2; a 14K-token Gemma E4B prompt went from 770 s to 631 s with identical output.
- Faster long-context generation for Gemma 4 12B and Qwen 3.8 Flash Next. Their decode attention now uses the barrier-free kernel the other models already had: at 110K tokens, 40 ms to 12 ms per layer on the 12B and 46 ms to 6 ms on Flash Next.
- Stricter local server. Like the OpenAI API, an unrecognized request field now returns a 400
unknown_parameternaming it instead of being silently ignored, with a suggestion for near misses (max_token) and a pointer from other servers' fields (chat_template_kwargstoenable_thinking). Unsupported OpenAI parameters and JSON response formats returnunsupported_value.reasoning_effortkeeps working for GPT-OSS. - Fixed: the
tuff promptandtuff servecommands bundled inside TUFF.app crashed with "unable to find bundle named TUFF_TUFFEngine" on Macs without a source checkout. - Long prefills set
AGX_RELAX_CDM_CTXSTORE_TIMEOUT=1so macOS is less likely to kill them while the display is active; set it yourself to override. GPU failures now name the layer and phase that failed.
Thanks to drumih/turbo-fieldfare, whose #159, #171 and #182 these changes build on.
Validation: all 1,640 tests passed locally. Every model shape matched a CPU reference, Gemma E4B and 12B produced token-identical output to 5.0.3, and the release archive passed extraction, code-signature, bundled launcher, checksum, and signed update-feed checks, with the packaged app, tuff prompt and tuff serve run from outside the repository.
Download the macOS arm64 ZIP below, or update through TUFF's built-in updater. The app is ad-hoc signed, not notarized.
TUFF v5.0.3
TUFF 5.0.3 restores the macOS Tahoe Liquid Glass appearance.
- 5.0.2 was built with Swift 6.4, whose SwiftPM links through clang with
--sysrootonly. The app was stamped as built against the macOS 15 SDK, so macOS ran it with the legacy appearance. Packaging now passes-isysrootto the link and refuses to produce an app stamped with an SDK older than 26. - The sidebar keeps the translucency from 5.0.2: the floating glass sidebar now blurs the desktop behind the window. It stays opaque with Reduce Transparency on.
Validation: all 1,610 tests passed locally. The release archive passed extraction, code-signature, bundled launcher, checksum, and signed update-feed checks, and the packaged app was compared side by side against 5.0.2.
Download the macOS arm64 ZIP below, or update through TUFF's built-in updater. The app is ad-hoc signed, not notarized.
TUFF v5.0.2
TUFF 5.0.2 fixes incomplete streamed responses and Swift 6.4 build errors.
- Preserve every incremental text delta when decode-service snapshots are throttled, including pending text before completion, cancellation, or failure (#2).
- Make the app and server decode callbacks compatible with Swift 6.4 region isolation while preserving sequential per-request state (#3).
- Accept both flat and structured SwiftPM font resource bundles during packaging.
- Add a transport-level regression test covering all three terminal events, leading whitespace, and Unicode text. The test reproduces the text loss on the previous implementation.
Thanks to @Andresit0 and @Osiriss664 for the fixes.
Validation: all 1,604 tests in 260 suites passed in final GitHub CI, including the new streaming regression's three terminal cases. The release archive passed extraction, code-signature, bundled launcher, font lookup, checksum, and signed update-feed checks.
Download the macOS arm64 ZIP below, or update through TUFF's built-in updater. The app is ad-hoc signed, not notarized.