Skip to content

vMLX 1.6.46

Choose a tag to compare

@jjang-ai jjang-ai released this 30 Aug 12:01

vMLX 1.6.46

Changes

  • feat(glm5_next): vendored GLM-5.3-Flash text runtime, registry row, text-route (d26ef226c)
  • feat(glm5_next): register the bundle-stamped parser ids as real aliases (8c6ae91b0)
  • fix(glm5_next): prefix caching fails closed until the typed native schema lands (f8b6f04e8)
  • fix(glm5_next): arm the family cache gate — runtime-type detection never fired (c79f8a2a4)
  • feat(glm5_next): DSA sparse indexer path — full context unlocked (df74393ab)
  • feat(panel): glm5-next family row, parser dropdown entries, alias canonicalization (5c5762ad6)
  • feat(panel): visual wired-limit recommendation popup at session launch (W0-W2 core) (8dda12ccf)
  • test: update engine-audit pin to depth-free adaptive MTP launcher (2b0992cf5)
  • fix(native-mtp): count glm5_next MTP block so health reports it honestly (T0) (81a22ab93)
  • Wire native MTP for GLM-5.3 (fc71d3883)
  • Activate GLM native MTP at the public loader (963872b64)
  • Map GLM MTP quantization onto the attached head (ffb4f2169)
  • Make text MTP depth policy actually adaptive (85d8b521d)
  • Warm GLM Metal kernels before readiness (4150ac442)
  • Group GLM affine projections for fewer dispatches (8cfc4166c)
  • Add peak benchmark profiles to the server panel (192b6b7e4)
  • Group Qwen4 decode projections at load time (8f3c537d7)
  • Record request-exact MTP benchmark telemetry (24099d504)
  • Fuse small-row gated RMSNorm for hybrid decode (0e58a3570)
  • Parallelize Qwen4 PLE decode row reads (cf93aaacd)
  • Add source-gated affine MoE pair fusion (144bf6b3b)
  • Cache completed GLM DSA pool keys (000db39dc)
  • Fuse Qwen GDN decode convolution state update (2748d3259)
  • Fuse GLM KDA decode convolutions (f3accebf6)
  • Fuse GLM KDA decode state update (247e0f319)
  • Fuse GLM mHC decode transform (1abb7cdcb)
  • Fuse Qwen PLE decode convolution (81787abcc)
  • Prefer full Qwen affine MoE decode fusion (198e7a284)
  • Fuse GLM affine MoE down reduction (379ee4cdf)
  • Fuse sparse index decode scoring (ced00f719)
  • Fuse GLM hyper-connection placement (d1fdcb940)
  • Standardize cross-family acceleration status (ecb25fe3b)
  • Report observed fused decode paths (be3f212f2)
  • Group Qwen3.5 GDN decode projections exactly (6741a0ec1)
  • Flush final MLLM detokenizer bytes to streams (80879777f)
  • Reconcile terminal MLLM stream suffixes (5dbd1780a)
  • Handle normalized terminal MLLM suffixes (d851fdc24)
  • Reconcile terminal visible stream suffixes (d327e881b)
  • Back off losing adaptive MTP probes (1acbba0d3)
  • Separate adaptive MTP seed from depth ceiling (095c003eb)
  • Warm adaptive MTP before probing depth (75112e52e)
  • Enable proven Qwen4 affine MoE pair decode (45d0bda76)
  • Preserve GLM mixed caches at MTP finalization (bae76a11f)
  • Keep unmeasured adaptive text MTP on AR (eabe3881b)
  • Enable proven GLM mHC decode fusion (d27d660ed)
  • Add opt-in Qwen3.5 GDN decode fusion (629e22498)
  • Preserve exact Qwen GDN activation math (6e785c275)
  • Keep Qwen GDN activation paths family-specific (5efa5bd5f)
  • Add exact Qwen3.5 GDN gate preparation candidate (b22745d20)
  • Expose DSV4 indexer acceleration status (cfc872ae4)
  • Separate DSV4 indexer hits from fallbacks (1d0fe00b2)
  • Add exact typed GLM native-state prefix caching (6feb91e0a)
  • Preserve GLM prompt cache through native MTP (037c0de25)
  • Bound GLM native cache snapshot memory (14a9a18d3)
  • Persist GLM prompt cache before final-token allocation (995d948f5)
  • Validate GLM typed cache metadata on live restore (7159d1cc2)
  • Enable GLM typed prompt cache in the app (00c4ce3e5)
  • Grade native prompt cache and MTP UI truth (e7e54af52)
  • Apply Native MTP sampling at chat request boundary (4d9304a54)
  • Accept native MTP and bundle alias proof truth (edfdd3d5e)
  • Make native MTP depth sweeps truly fixed (4953fee16)
  • Reuse matched AR controls in MTP depth sweeps (cc1ef0904)
  • Reject native MTP depths slower than AR (68dbdab11)
  • Fix fast agentic replay evidence capture (eb5caedc2)
  • Observe gateway aborts on bound backend (f2f12789b)
  • Keep busy request lifecycle health live (92e059d24)
  • Align agentic release proof with transmitted requests (7db711a58)
  • Fix release proof grading for current UI and cache telemetry (adbb49b48)
  • Read cache percentage from recorded engine command (ca63de0e4)
  • Prepare vMLX 1.6.46 (ef609b26f)
  • Update typed-cache UI audit for generalized policy (5a953522a)

Runtime provenance

  • vMLX source: 5a953522a61e3d2b9458c480e8724a9da5066811
  • JANG runtime: 54b29d29916fefa60ce15dce72e3735757f996eb (jang 2.5.47)
  • Tahoe/macOS 26 is the default build; Sequoia is the compatibility build.
  • Both DMGs are Developer ID signed, independently notarized, stapled, and Gatekeeper checked from the final release handoff.

Downloads

  • Tahoe: vMLX-1.6.46-tahoe-arm64.dmg
    • SHA-256: 524174decd435b47d37e17653d489438aab14357052baf634d5e3e59d64cb676
    • Size: 538124488 bytes
  • Sequoia: vMLX-1.6.46-sequoia-arm64.dmg
    • SHA-256: 6e5a0e23297bb55b2bf2d41690d6e004874b5a781d655bd9f5f82f0d85b05ac5
    • Size: 515655536 bytes

Python distributions

  • Wheel: vmlx-1.6.46-py3-none-any.whl — SHA-256 56c08c8e4054a712f5434279ce9eebe36ce29b3be0bac3129812ecea1347ffe8
  • Source: vmlx-1.6.46.tar.gz — SHA-256 a95020dba01e99e50a51549fa41306c8355339fde8a57284a98ff2a6bbe0a51a