Repository navigation
TUFF 7.1.0
TUFF 7.1.0 keeps one local server, the model-routing Background API introduced in 7.0.0, and makes it work with coding agents such as oh-my-pi.
- One server. The Background API and
tuff serverun the same server: every installed model on one loopback endpoint, loaded when a request names it and unloaded when idle.defaultmeans your chosen model. Each model runs with its catalog context, expert-cache and prefill settings for your Mac. - Server screen. Rebuilt around the Background API: an on/off switch, the endpoint, default model, unload delay and port, plus the loaded model with its unload countdown, active and queued requests, and Open Log. Clone builds show the command to run the same server in Terminal.
- Qwen with oh-my-pi. Qwen clients send
preserve_thinkingandchat_template_kwargs, which the server used to refuse with a 400. It now acceptschat_template_kwargsholdingenable_thinkingorpreserve_thinkingand still refuses any other template argument.
Removed: the app's Start/Stop local server, which served only Chat's loaded model, and fixed-model serving with tuff serve --model / TUFFServer --model and its tuning flags. A removed flag prints what to use instead. --all-models is still accepted and has no effect.
Known limits: oh-my-pi's full default toolset fails on Gemma models because its task tool uses a schema Gemma's tool template cannot represent; after a Qwen tool call the next request re-reads the whole prompt; Qwen3.8 Flash Next's 2,048-token context only just fits oh-my-pi's prompt without tools.
Inference kernels and model packs are unchanged from 7.0.0.
Validation on one 16 GB M2 MacBook Air: 1,698 Swift tests and the model-free gate, packaging and isolated updater fixtures; the packaged CLI; and oh-my-pi through the packaged server on Gemma 4 E4B and 26B-A4B, Qwen3.6 and Qwen3.8 Flash Next, including tool calls. Other Macs were not tested. See the release validation for details and limits.
The app is arm64 and ad-hoc signed, not notarized. The release includes the ZIP, a SHA-256 checksum and a signed Sparkle update feed.
AI assistance: Claude implemented, tested and documented these changes, and OpenAI Codex reviewed them.