Skip to content

vMLX 1.6.29

Choose a tag to compare

@jjang-ai jjang-ai released this 15 Aug 10:40
· 791 commits to main since this release

vMLX 1.6.29

  • Hybrid prefix cache: follow-up turns on long documents up to 20x faster (companion state completes paged-rebuilt bases; 43.7k-token doc follow-ups 107s -> 5s, byte-identical at temperature 0).
  • 100k-context multiturn proven stable: 97.6k-token conversations retrieve planted facts with 99.9 percent cache reuse and no memory faults.
  • Chunked SSM re-derive ships default-on: long-context reuse restored from 0 to 99.9 percent above 12.5k tokens on hybrid families.
  • Contentless assistant turns no longer 500 on either chat completions or the Responses API.
  • Per-family cache, parser, and generation-default hardening verified across 12 model families, including DSV4 Flash native composite caching, Gemma 4 mixed-SWA with audio, Qwen 3.6/3.8 hybrid lines with native MTP, and TurboQuant KV bundles.
  • Release integrity: version stamps unified across five surfaces; both DMGs notarized and stapled.