Skip to content

vMLX 1.6.15

Choose a tag to compare

@jjang-ai jjang-ai released this 22 Jul 09:05

vMLX 1.6.15 is a reliability and runtime checkpoint for local Apple-silicon inference.

Highlights:

  • Separates reasoning, visible content, and tool-call streaming across Chat Completions, Responses, Anthropic, Ollama, and Electron chat paths.
  • Hardens post-tool continuation and removes speculative or phantom tool-start emissions.
  • Improves Laguna S2.1 reasoning/template behavior and q4 TurboQuant mixed-SWA cache restore, including disk-only partial-prefix reuse.
  • Preserves model-derived sampling and parser defaults across saved sessions.
  • Improves single-model gateway swaps, disconnect handling, and process cleanup.
  • Adds cache/runtime fixes for LFM, MiniMax M3, Nemotron Omni, Gemma 4 media, Qwen MTP video, and DSV4 long prefill.
  • Fixes release packaging of bundled Python native components.

The signed Sequoia and Tahoe applications are distributed from the companion MLXStudio release.

Thanks to @hornsan1 for continued testing and feedback.