vMLX 1.6.15
vMLX 1.6.15 is a reliability and runtime checkpoint for local Apple-silicon inference.
Highlights:
- Separates reasoning, visible content, and tool-call streaming across Chat Completions, Responses, Anthropic, Ollama, and Electron chat paths.
- Hardens post-tool continuation and removes speculative or phantom tool-start emissions.
- Improves Laguna S2.1 reasoning/template behavior and q4 TurboQuant mixed-SWA cache restore, including disk-only partial-prefix reuse.
- Preserves model-derived sampling and parser defaults across saved sessions.
- Improves single-model gateway swaps, disconnect handling, and process cleanup.
- Adds cache/runtime fixes for LFM, MiniMax M3, Nemotron Omni, Gemma 4 media, Qwen MTP video, and DSV4 long prefill.
- Fixes release packaging of bundled Python native components.
The signed Sequoia and Tahoe applications are distributed from the companion MLXStudio release.
Thanks to @hornsan1 for continued testing and feedback.