Skip to content

TUFF 7.1.0

Choose a tag to compare

@rexmhall09 rexmhall09 released this 03 Oct 05:42
· 3 commits to main since this release

TUFF 7.1.0 keeps one local server, the model-routing Background API introduced in 7.0.0, and makes it work with coding agents such as oh-my-pi.

  • One server. The Background API and tuff serve run the same server: every installed model on one loopback endpoint, loaded when a request names it and unloaded when idle. default means your chosen model. Each model runs with its catalog context, expert-cache and prefill settings for your Mac.
  • Server screen. Rebuilt around the Background API: an on/off switch, the endpoint, default model, unload delay and port, plus the loaded model with its unload countdown, active and queued requests, and Open Log. Clone builds show the command to run the same server in Terminal.
  • Qwen with oh-my-pi. Qwen clients send preserve_thinking and chat_template_kwargs, which the server used to refuse with a 400. It now accepts chat_template_kwargs holding enable_thinking or preserve_thinking and still refuses any other template argument.

Removed: the app's Start/Stop local server, which served only Chat's loaded model, and fixed-model serving with tuff serve --model / TUFFServer --model and its tuning flags. A removed flag prints what to use instead. --all-models is still accepted and has no effect.

Known limits: oh-my-pi's full default toolset fails on Gemma models because its task tool uses a schema Gemma's tool template cannot represent; after a Qwen tool call the next request re-reads the whole prompt; Qwen3.8 Flash Next's 2,048-token context only just fits oh-my-pi's prompt without tools.

Inference kernels and model packs are unchanged from 7.0.0.

Validation on one 16 GB M2 MacBook Air: 1,698 Swift tests and the model-free gate, packaging and isolated updater fixtures; the packaged CLI; and oh-my-pi through the packaged server on Gemma 4 E4B and 26B-A4B, Qwen3.6 and Qwen3.8 Flash Next, including tool calls. Other Macs were not tested. See the release validation for details and limits.

The app is arm64 and ad-hoc signed, not notarized. The release includes the ZIP, a SHA-256 checksum and a signed Sparkle update feed.

AI assistance: Claude implemented, tested and documented these changes, and OpenAI Codex reviewed them.