Skip to content

3.32.3

Latest

Choose a tag to compare

@davidkoski davidkoski released this 30 Sep 20:38
· 29 commits to main since this release
3b339ad
  • pick up mlx-swift 0.32.3
  • massive improves in draft models, KVCache handling, tool handling
  • add MLXFoundationModels and MLXGuidedGeneration
  • bug fixes

What's Changed

  • Set indentConditionalCompilationBlocks to false in .swift-format by @ctymoszek in #385
  • CI: fix mac_build_and_test for mlx-swift >= 0.31.5 (skip package plugin validation) + swift-format 603 conformance by @GoodOlClint in #404
  • ci: pin swift-format version so lint doesn't drift across PRs by @GoodOlClint in #386
  • Gemma 3: chunked prompt prefill, skip lm_head on prompt positions by @beshkenadze in #346
  • Fix Gemma 4 VLM load: KV-shared layers must not declare k_proj/v_proj by @fdagostino in #390
  • add note on swift-format version and fill out configuration by @davidkoski in #416
  • ChatSession: cancel the generation task when the consumer goes away by @fdagostino in #413
  • Qwen3 embedder: honor attentionMask instead of silently ignoring it by @spokvulcan in #418
  • Qwen3VL: apply the sRGB tone curve in the image preprocess path by @spokvulcan in #411
  • Convert Gemma tool parameters by schema type by @atdrendel in #388
  • Fix prefill LMOutput.State being dropped on TokenIterator's .logits path by @NivDvir in #419
  • Fix Qwen3.5 VLM garbage output: gate the RMSNorm weight shift in sanitize by @GoodOlClint in #403
  • fix: Honor safetensors index when loading weights by @Zrxrxrx in #408
  • Qwen3.5/3.6: windowed prefill + state-threaded warm continuation (fixes multi-turn M-RoPE drift) by @spokvulcan in #399
  • Make DeepseekV3Model.init public by @magicnight in #422
  • Make MLXLLM compilable on Linux by @gmondada in #321
  • Fix tool round-trip: record assistant tool_calls message before tool results by @jacobwillemsma in #409
  • Add cooperative cancellation check to the prompt prefill loop by @sunshinfight in #423
  • Fix cancellation submission race by @SpiraMira in #389
  • Add MLXFoundationModels: an MLX-backed FoundationModels LanguageModel by @ctymoszek in #334
  • Update project name in CONTRIBUTING.md by @CharlieTLe in #427
  • Decode PoolingConfiguration with optional pooling_mode flags by @GoodOlClint in #376
  • Qwen3VL: default per-image resolution to a 1,280 vision-token budget by @spokvulcan in #398
  • Fix silent state truncation for non-multiple-of-32 GDN key dim by @GoodOlClint in #381
  • Document model compatibility requirements by @RNT56 in #303
  • Fix MLXVLM Gemma4 loader to honor num_kv_shared_layers (E-series targets) by @GoodOlClint in #384
  • Document -skipPackagePluginValidation for local test runs by @CharlieTLe in #430
  • Fix Python finfo.min port in quantized attention masking by @ronaldmannak in #369
  • Implement Gemma 4 E-series MTP centroid embedder (use_ordered_embeddings) by @GoodOlClint in #383
  • Update link to CONTRIBUTING document in PR template by @CharlieTLe in #428
  • Gemma 4: aspect-preserving image resize with per-image dynamic token counts; fix multi-image requests by @prbme in #405
  • Track the current SDK's SamplingMode.Kind case names by @CharlieTLe in #431
  • update tag on swift-format -- I had it wrong by @davidkoski in #421
  • Fix FoundationModels API drift and the integration tests that no longer compiled by @thechriswebb in #438
  • Prepare tool-calling input through the model's processor (fixes #433) by @jyauxi in #435
  • Make BaseKVCache.ropeOffset overridable by @GoodOlClint in #437
  • fix(MLXFoundationModels): stop respond() crashing when emitting usage on the FM-27 SDK by @thechriswebb in #439
  • Silence Metal compiler warning noise in integration test runs by @CharlieTLe in #440
  • add TurboQuant KV cache compression by @TheTom in #232
  • Expose minimal Gemma3 surface for frozen-text-encoder use by @xocialize in #387
  • Load EOS token IDs from nested text configs by @aleroot in #449
  • Optimize Qwen3.5 interleaved M-RoPE by @aleroot in #442
  • MLXLLM Gemma4Text: add MoE block (router + experts) for the text path by @neuromechanist in #364
  • Fix MTP integration test skip semantics; auto-download 31B checkpoint pair by @CharlieTLe in #445
  • Enable MTP speculative decoding for Gemma 4 12B (gemma4_unified target + gemma4_unified_assistant drafter) by @fdagostino in #415
  • Hoist tool-schema $defs to the tool-calling envelope root (fixes #432) by @jyauxi in #434
  • Qwen3VL vision: stop using a huge amount of memory to process images. by @kklimuk in #455
  • Add regression coverage for mixed-precision (per-layer) quantized checkpoint loading by @fdagostino in #395
  • DeepSeek-V3: populate kvHeads and update the KV cache once per attention call by @GoodOlClint in #457
  • Support multi-round tool calling in MLXFoundationModels by @thechriswebb in #456
  • Add DeepSeek-V2 model by @GoodOlClint in #379
  • Add Hunyuan dense V1 (hunyuan_v1_dense): Hunyuan-MT-7B and Hy-MT2-7B by @beshkenadze in #347
  • Add GitHub workflow for IntegrationTesting tests by @CharlieTLe in #458
  • Add end-to-end video input to the base Gemma 4 (gemma4) VLM by @fdagostino in #391
  • Integration tests: build on both macOS 26 and 27 SDKs by @CharlieTLe in #464
  • Add routing-correctness tests for GLM4MOE / GLM4MOELite MoE gates by @GoodOlClint in #463
  • Fix guided generation dropping required properties (streaming detokenizer) by @thechriswebb in #465
  • Integration tests CI: run on Xcode 26 only by @CharlieTLe in #477
  • Gate remaining FoundationModels integration tests behind the 27 SDK by @CharlieTLe in #478
  • Integration tests: scope down heavy suites, fix flakes and memory pressure by @CharlieTLe in #480
  • Add Nanbeige4.2 (looped transformer) model support by @spokvulcan in #460
  • Qwen3.5/3.6: run the decode step through compiled traces by @spokvulcan in #467
  • Qwen3.5/3.6 MoE: fused router top-k kernel for decode (bit-identical, CI-pinned) by @spokvulcan in #469
  • Fix Gemma 3n Boolean attention masks by @codesworth in #479
  • Fix lfm2 toolcall parser and glm4 gate GLM-4-0414 by @CharlieTLe in #481
  • Qwen3.5/3.6: fold the GDN decode conv into the compiled step by @spokvulcan in #468
  • Fix structured ChatSession continuations by @aleroot in #472
  • Let models declare chat conventions via ChatConventionsProviding by @CharlieTLe in #482
  • Qwen2.5-VL / Qwen2-VL: windowed prefill + state-threaded warm continuation (wires #420) by @NivDvir in #448
  • Fix remaining Linux build breaks in MLXLMCommon and MLXGuidedGeneration by @GoodOlClint in #483
  • Fix gated delta recurrence precision by @aleroot in #488
  • Fix TF32 test-oracle flakiness in the Gemma4 chunk-invariance tests by @GoodOlClint in #485
  • Add support for Qwen3-VL-MoE models by @lucasnewman in #322
  • Migrate all models to ChatConventionsProviding; delete the infer chains by @CharlieTLe in #502
  • Add custom LogitProcessor injection to GenerateParameters / ChatSession by @xocialize in #401
  • Run Integration Tests nightly and report failures on a tracking issue by @CharlieTLe in #513
  • Stand down from MTP speculation before the sliding cache wraps by @angelsbrood in #506
  • Add typed KV cache configuration and runtime reporting by @aleroot in #453
  • Add parser for GPT-OSS Harmony tool call format by @aleroot in #146
  • Prefill: balanced chunking behind a new PrefillParameters (~9% at 32K) by @spokvulcan in #470
  • Olmo3: fix newCache signature so the sliding-window cache is actually used by @GoodOlClint in #462
  • Prompt cache: persist model state, wire Qwen3-VL and GLM-OCR restoration, and fail closed by @kklimuk in #475
  • Add single-dispatch TurboFlash for short contexts by @aleroot in #520
  • Update Package.swift by @ActuallyTaylor in #519
  • Fix tied quantized Qwen3 MoE head sanitization by @JamesPriceZV in #490
  • Speculate past the sliding window in MTP by @angelsbrood in #516
  • Remove the untyped staged-cache fallback by @angelsbrood in #526
  • Add Muse-Glimmer, a 30B agentic multimodal model, with ATEM tool calling by @CharlieTLe in #523
  • Enforce and expose effective maxKVSize cache policy by @aleroot in #514
  • Add protocol-aware thinking budget enforcement by @aleroot in #521
  • Add Qwen3.5 MTP speculative decoding by @aleroot in #351
  • Add Qwen 3.5 JSON tool-call fallback by @aleroot in #529
  • bump mlx-swift to 0.31.6 -- .4 was missing new API that the current code calls by @davidkoski in #484
  • Expose Gemma 4 encoder surface at @_spi(GemmaEncoder) by @xocialize in #530
  • Enforce copy semantics on LogitProcessor for speculative decoding (#508) by @aleroot in #533
  • Fix tied word embedding weight sanitization by @aleroot in #532
  • Add TranslateGemma support (Gemma 3 translation template) by @beshkenadze in #348
  • Support video processing options in UserInput and MediaProcessing by @aleroot in #534
  • Surface rejected tool calls as generation events by @aleroot in #512
  • Add an Xcode 27 compile check and handle Generation.rejectedToolCall by @CharlieTLe in #538
  • Fix nested tool grammar schema test by @angelsbrood in #527
  • Qwen2.5-VL: accumulate vision cuSeqlens across temporal slices by @NivDvir in #509
  • Add Qwen3.5 VLM processor config fallback by @ngutech21 in #546
  • Tool call parser hardening by @aleroot in #531
  • Muse Glimmer: honor the requested KV-cache capacity by @aleroot in #556
  • Report prompt-cache reuse in GenerateCompletionInfo by @aleroot in #559
  • Fix Qwen3VL prompt cache reuse for text-only inputs by @ngutech21 in #549
  • Add opt-in q4_0 lattice calibration to model conversion by @GoodOlClint in #507
  • Add Helium (Kyutai) LLM port by @noorbhatia in #555
  • Read Gemma's brace-form object and array values by @aleroot in #557
  • Add LoRA dropout and training mode support by @aleroot in #541
  • Add reranker API support by @aleroot in #375
  • Load weight files a safetensors index leaves out by @aleroot in #562
  • Fix MLXFoundationModels compilation against the Xcode 27 beta 5 SDK by @abra-code in #544
  • Tune TurboFlash query grouping for short decode by @aleroot in #570
  • Fix mixed-precision QLoRA output promotion by @aleroot in #564
  • Add extensible VLM processor loading rules by @aleroot in #565
  • Share fused Qwen3.5 router top-k across LLM and VLM by @ngutech21 in #567
  • Fix SSM gradient transforms across chunks by @aleroot in #571
  • MLXVLM: resolve the LFM2VL image token id from the vocabulary by @Siddhesh2377 in #576
  • Strengthen chunked SSM gradient tests by @aleroot in #574
  • Parallelize weight loading with byte-balanced contiguous groups(up to ~1.8x faster model loads) by @aleroot in #575
  • Add variance-normalized KV cache by @aleroot in #329
  • Temporarily disable the Xcode 27 integration build on pull requests by @thechriswebb in #582
  • Replace MLXRandom.seed with task-local withRandomState in tests by @aleroot in #583
  • Add ai usage policy by @ctymoszek in #585
  • Generalise compiled decode segments to Qwen 3 Next by @aleroot in #569
  • Share fused MoE router top-k across models by @aleroot in #568
  • Qwen direct reduction by @aleroot in #573
  • Fuse Qwen GDN input projections by @aleroot in #572
  • Clear the compiler warnings from the library and both test targets by @thechriswebb in #594
  • Leverage GitHub workflows to keep PR labels up-to-date by @thechriswebb in #587
  • Qwen3.5: don't shift norm weights just because a checkpoint has MTP tensors by @mandrael in #598
  • Fix DeepSeek MTP layer filtering by @aleroot in #592
  • Allow downstream specialization of the Qwen3.5 GDN/MoE blocks and SwitchGLU by @agerjura in #511
  • Expose Falcon-H1 encoder surface at @_spi(FalconH1Encoder) by @xocialize in #596
  • Report prompt token counts to FoundationModels again by @thechriswebb in #591
  • Suspend instead of blocking cooperative threads while weights load by @aleroot in #579
  • Harden the Foundation Models test fixtures by @thechriswebb in #599
  • ParoQuant: extend to MoE and make the whole path fast (#164 follow-up) by @spokvulcan in #471
  • Await the async loadWeights overload in the IntegrationTesting tests by @thechriswebb in #605
  • Declare module weights as compile state in compiled decode paths by @aleroot in #589
  • Gemma4Text: expose decoder layers as loraLayers so MLP LoRA targets load by @cives93 in #602
  • Let training differentiate the gated-delta recurrence by @jhancock1975 in #616
  • Emit only the new scalars when a token extends the previous character by @spokvulcan in #613
  • Gemma 4 VLM: fuse float32 logit softcapping by @aleroot in #615
  • Fix generation loop blocking Swift cooperative workers by @aleroot in #611
  • Add configuration-based LoRA metadata discovery by @aleroot in #597
  • RotatingKVCache: make trim wrap-aware instead of corrupting the ring by @aleroot in #584
  • Add bounded cross-dialect tool-call recovery by @aleroot in #548
  • ChatSession: reuse the KV cache for append-only media turns by @NivDvir in #515
  • Clear the MLX cache on the first generated token by @negativetime in #620
  • Extract the model cache and stop whole-cache eviction abandoning an in-flight load by @thechriswebb in #603
  • pick up mlx-swift 0.32.2 by @davidkoski in #646
  • Named image attachments in MLXLMCommon and the Foundation Models adapter by @thechriswebb in #535
  • Convert the booleans in Foundation Models tool schemas to Swift Bool values by @thechriswebb in #653
  • be able to trigger workflows on main. notes about upcoming releases. by @davidkoski in #647
  • require mlx-swift-0.32.3 for SDK issue by @davidkoski in #657
  • adjustments for running tests on NAX/TF32 hardware by @davidkoski in #649

New Contributors

Full Changelog: 3.31.4...3.32.3