Skip to content

vMLX 1.6.54

Latest

Choose a tag to compare

@jjang-ai jjang-ai released this 06 Sep 06:08
· 25 commits to main since this release

vMLX 1.6.54

Changes

  • fix(cache): assert the hybrid restore invariant — KV offset must equal the companion boundary (0c4f47c6)
  • docs: note the hybrid restore alignment check in the 1.6.53 changelog (a662d86a7)
  • Publish vMLX 1.6.53 updater manifest (6ce62fb1)
  • perf(mtp): depth ladder D3 -> D1 -> AR, margin 1.0, cheap probes, promotion back to D3 (47000bee)
  • perf(mtp): start at D1 after a request that ended in AR or D1; no context scale on measured baselines (1743c2be)
  • chore(mtp): log AR handoff and re-entry seed wall time (fe4c6e86)
  • perf(mtp): per-request retry budget for re-entry and promotion probes (7dd5a2df)
  • Prepare vMLX 1.6.54 (e928f7e7)
  • fix(tools): accept JSON-schema native tool prompts for MiniMax-M2.7, Qwen3.8-27B and Nanbeige (e4f4656d)
  • fix(panel): resolve npm for coding-tool installs; Hermes manual setup snippet (5097fd1a)
  • docs: 1.6.54 changelog for the tool-prompt and coding-tool install fixes (888dfdda)
  • test: pin both native schema dialects in the xml_function tool-prompt rule (3381db58)
  • qwen4_exp: retain completed QSA index pools across decode calls (1b78268e)
  • qwen4_exp: opt-in multi-row gates for grouped GDN projections and compiled HC (cdac72ba)
  • qwen4_exp: build the QSA selection mask integers on device (5532fe87)
  • native MTP governor: seed-uncertainty margin, no sticky start from a cold request (4ecbb8aa)
  • changelog: 1.6.54 QSA pool retention, exact index angles, seed-margin governor (f04fabcf)
  • qwen4_exp PLE: copy each distinct row once in the mmap gather (a1d8e341)
  • qwen4_exp: MoE route-overlap diagnostic for multi-row verification (d623cf2b)
  • sparse index cache: grow the raw index lane in steps instead of per-token copies (bd48624d)
  • qwen4_exp: group GDN projections and compile HC at MTP verify widths by default (33fd97f1)
  • native MTP: measured depth economics probe (configured depth vs D1) (f0b3ebdf)
  • changelog: 1.6.54 verify-width grouping, index step growth, depth economics, PLE dedupe (e63a8033)
  • affine MoE pair kernel: opt-in mixed gate/up layouts and 3/6-bit experts (23a3b92b)
  • sparse index cache: validate the initialized extent, count retained bytes, gate the rotary change (8e8737c5)
  • qwen4_exp: qualify the exact rotary angle separately; opt-in exact angle for attention (f61d94a9)
  • affine MoE pair kernel: keep the native-MTP draft head on the stock path (04b4786d)
  • changelog: 1.6.54 pair-kernel head exclusion, mixed-layout kernel, index validation, rotary gate (ba5b0d45)
  • affine MoE pair kernel: stay on the stock path while native MTP is active (19c8ef35)
  • affine MoE pair kernel: mixed gate/up and 3/6-bit layouts on by default for qwen4_exp (43d1707b)
  • changelog: 1.6.54 pair kernel off under native MTP; mixed layouts default on (b12588ce)
  • qwen4_exp: verify-width grouping on by default, with the full-model receipts (3fe8b5d6)
  • native MTP: state the fixed-vs-governed depth contract; report depth occupancy (0674ceb5)
  • native MTP: fixed keeps its depth except through the explicit AR-safety valve (aef08444)
  • native MTP: per-cycle trace switch for governor audits (cf94772e)
  • native MTP: bounded AR calibration for a context-matched baseline (3daf176e)
  • Native MTP: describe the bounded AR calibration in the 1.6.54 changelog (880287f2)
  • Native MTP governor: calibration resumes the run, keeps its own budget, never teaches a loss (5d1339c3)
  • changelog: 1.6.54 native MTP calibration entry describes the resumed-run semantics and the measured governor cost (0b5ef741)
  • Native MTP governor: no re-entry cap, geometric backoff with relaxation, confirmed final rung (072a82ea)
  • changelog: 1.6.54 governor entry covers the uncapped re-entry backoff and the confirmed final rung (a0c98903)
  • Tool calls: an explicit null for a required nullable argument is a value, not a missing argument (40f81a1d)
  • Release build: signature checks read the codesign output through here-strings, not printf pipes (14444461)
  • Native MTP governor: calibration cost bounded by spacing, stall-time readings re-checked early (cb042a4d)

Runtime provenance

  • vMLX source: cb042a4d2e45936faafa2c3645d328636056f592
  • JANG runtime: f3c79081c3d99e8a26fa28444934f7567acb00c6 (jang 2.5.47)
  • Tahoe/macOS 26 is the default build; Sequoia is the compatibility build.
  • Both DMGs are Developer ID signed, independently notarized, stapled, and Gatekeeper checked by the candidate workflow.

Downloads

  • Tahoe: vMLX-1.6.54-tahoe-arm64.dmg
    • SHA-256: 0b0b46e2ae34288324371c0cb243c9f9dd17f366d85e9a30f85f84dd9c4703fb
    • Size: 556364115 bytes
  • Sequoia: vMLX-1.6.54-sequoia-arm64.dmg
    • SHA-256: 46f84a6b37af77a7ad31ed5781769839b32a606bec7364f43112d56838442802
    • Size: 534190592 bytes

Python distributions

  • Wheel: vmlx-1.6.54-py3-none-any.whl — SHA-256 de9dca3a782fe55ab3dd284e893b76f0e64a0a5c45bd7813a1577795db8f10e2
  • Source: vmlx-1.6.54.tar.gz — SHA-256 20f1f37cb5304d09e5f6a265ebe5a5a73cfdf302da4ca2cb265a84412d77f29c