Repository navigation
- pick up mlx-swift 0.32.3
- massive improves in draft models, KVCache handling, tool handling
- add MLXFoundationModels and MLXGuidedGeneration
- bug fixes
What's Changed
- Set indentConditionalCompilationBlocks to false in .swift-format by @ctymoszek in #385
- CI: fix mac_build_and_test for mlx-swift >= 0.31.5 (skip package plugin validation) + swift-format 603 conformance by @GoodOlClint in #404
- ci: pin swift-format version so lint doesn't drift across PRs by @GoodOlClint in #386
- Gemma 3: chunked prompt prefill, skip lm_head on prompt positions by @beshkenadze in #346
- Fix Gemma 4 VLM load: KV-shared layers must not declare k_proj/v_proj by @fdagostino in #390
- add note on swift-format version and fill out configuration by @davidkoski in #416
- ChatSession: cancel the generation task when the consumer goes away by @fdagostino in #413
- Qwen3 embedder: honor attentionMask instead of silently ignoring it by @spokvulcan in #418
- Qwen3VL: apply the sRGB tone curve in the image preprocess path by @spokvulcan in #411
- Convert Gemma tool parameters by schema type by @atdrendel in #388
- Fix prefill LMOutput.State being dropped on TokenIterator's .logits path by @NivDvir in #419
- Fix Qwen3.5 VLM garbage output: gate the RMSNorm weight shift in sanitize by @GoodOlClint in #403
- fix: Honor safetensors index when loading weights by @Zrxrxrx in #408
- Qwen3.5/3.6: windowed prefill + state-threaded warm continuation (fixes multi-turn M-RoPE drift) by @spokvulcan in #399
- Make DeepseekV3Model.init public by @magicnight in #422
- Make MLXLLM compilable on Linux by @gmondada in #321
- Fix tool round-trip: record assistant tool_calls message before tool results by @jacobwillemsma in #409
- Add cooperative cancellation check to the prompt prefill loop by @sunshinfight in #423
- Fix cancellation submission race by @SpiraMira in #389
- Add MLXFoundationModels: an MLX-backed FoundationModels LanguageModel by @ctymoszek in #334
- Update project name in CONTRIBUTING.md by @CharlieTLe in #427
- Decode PoolingConfiguration with optional pooling_mode flags by @GoodOlClint in #376
- Qwen3VL: default per-image resolution to a 1,280 vision-token budget by @spokvulcan in #398
- Fix silent state truncation for non-multiple-of-32 GDN key dim by @GoodOlClint in #381
- Document model compatibility requirements by @RNT56 in #303
- Fix MLXVLM Gemma4 loader to honor num_kv_shared_layers (E-series targets) by @GoodOlClint in #384
- Document -skipPackagePluginValidation for local test runs by @CharlieTLe in #430
- Fix Python finfo.min port in quantized attention masking by @ronaldmannak in #369
- Implement Gemma 4 E-series MTP centroid embedder (use_ordered_embeddings) by @GoodOlClint in #383
- Update link to CONTRIBUTING document in PR template by @CharlieTLe in #428
- Gemma 4: aspect-preserving image resize with per-image dynamic token counts; fix multi-image requests by @prbme in #405
- Track the current SDK's SamplingMode.Kind case names by @CharlieTLe in #431
- update tag on swift-format -- I had it wrong by @davidkoski in #421
- Fix FoundationModels API drift and the integration tests that no longer compiled by @thechriswebb in #438
- Prepare tool-calling input through the model's processor (fixes #433) by @jyauxi in #435
- Make BaseKVCache.ropeOffset overridable by @GoodOlClint in #437
- fix(MLXFoundationModels): stop respond() crashing when emitting usage on the FM-27 SDK by @thechriswebb in #439
- Silence Metal compiler warning noise in integration test runs by @CharlieTLe in #440
- add TurboQuant KV cache compression by @TheTom in #232
- Expose minimal Gemma3 surface for frozen-text-encoder use by @xocialize in #387
- Load EOS token IDs from nested text configs by @aleroot in #449
- Optimize Qwen3.5 interleaved M-RoPE by @aleroot in #442
- MLXLLM Gemma4Text: add MoE block (router + experts) for the text path by @neuromechanist in #364
- Fix MTP integration test skip semantics; auto-download 31B checkpoint pair by @CharlieTLe in #445
- Enable MTP speculative decoding for Gemma 4 12B (gemma4_unified target + gemma4_unified_assistant drafter) by @fdagostino in #415
- Hoist tool-schema $defs to the tool-calling envelope root (fixes #432) by @jyauxi in #434
- Qwen3VL vision: stop using a huge amount of memory to process images. by @kklimuk in #455
- Add regression coverage for mixed-precision (per-layer) quantized checkpoint loading by @fdagostino in #395
- DeepSeek-V3: populate kvHeads and update the KV cache once per attention call by @GoodOlClint in #457
- Support multi-round tool calling in MLXFoundationModels by @thechriswebb in #456
- Add DeepSeek-V2 model by @GoodOlClint in #379
- Add Hunyuan dense V1 (hunyuan_v1_dense): Hunyuan-MT-7B and Hy-MT2-7B by @beshkenadze in #347
- Add GitHub workflow for IntegrationTesting tests by @CharlieTLe in #458
- Add end-to-end video input to the base Gemma 4 (gemma4) VLM by @fdagostino in #391
- Integration tests: build on both macOS 26 and 27 SDKs by @CharlieTLe in #464
- Add routing-correctness tests for GLM4MOE / GLM4MOELite MoE gates by @GoodOlClint in #463
- Fix guided generation dropping required properties (streaming detokenizer) by @thechriswebb in #465
- Integration tests CI: run on Xcode 26 only by @CharlieTLe in #477
- Gate remaining FoundationModels integration tests behind the 27 SDK by @CharlieTLe in #478
- Integration tests: scope down heavy suites, fix flakes and memory pressure by @CharlieTLe in #480
- Add Nanbeige4.2 (looped transformer) model support by @spokvulcan in #460
- Qwen3.5/3.6: run the decode step through compiled traces by @spokvulcan in #467
- Qwen3.5/3.6 MoE: fused router top-k kernel for decode (bit-identical, CI-pinned) by @spokvulcan in #469
- Fix Gemma 3n Boolean attention masks by @codesworth in #479
- Fix lfm2 toolcall parser and glm4 gate GLM-4-0414 by @CharlieTLe in #481
- Qwen3.5/3.6: fold the GDN decode conv into the compiled step by @spokvulcan in #468
- Fix structured ChatSession continuations by @aleroot in #472
- Let models declare chat conventions via ChatConventionsProviding by @CharlieTLe in #482
- Qwen2.5-VL / Qwen2-VL: windowed prefill + state-threaded warm continuation (wires #420) by @NivDvir in #448
- Fix remaining Linux build breaks in MLXLMCommon and MLXGuidedGeneration by @GoodOlClint in #483
- Fix gated delta recurrence precision by @aleroot in #488
- Fix TF32 test-oracle flakiness in the Gemma4 chunk-invariance tests by @GoodOlClint in #485
- Add support for Qwen3-VL-MoE models by @lucasnewman in #322
- Migrate all models to ChatConventionsProviding; delete the infer chains by @CharlieTLe in #502
- Add custom LogitProcessor injection to GenerateParameters / ChatSession by @xocialize in #401
- Run Integration Tests nightly and report failures on a tracking issue by @CharlieTLe in #513
- Stand down from MTP speculation before the sliding cache wraps by @angelsbrood in #506
- Add typed KV cache configuration and runtime reporting by @aleroot in #453
- Add parser for GPT-OSS Harmony tool call format by @aleroot in #146
- Prefill: balanced chunking behind a new PrefillParameters (~9% at 32K) by @spokvulcan in #470
- Olmo3: fix newCache signature so the sliding-window cache is actually used by @GoodOlClint in #462
- Prompt cache: persist model state, wire Qwen3-VL and GLM-OCR restoration, and fail closed by @kklimuk in #475
- Add single-dispatch TurboFlash for short contexts by @aleroot in #520
- Update Package.swift by @ActuallyTaylor in #519
- Fix tied quantized Qwen3 MoE head sanitization by @JamesPriceZV in #490
- Speculate past the sliding window in MTP by @angelsbrood in #516
- Remove the untyped staged-cache fallback by @angelsbrood in #526
- Add Muse-Glimmer, a 30B agentic multimodal model, with ATEM tool calling by @CharlieTLe in #523
- Enforce and expose effective maxKVSize cache policy by @aleroot in #514
- Add protocol-aware thinking budget enforcement by @aleroot in #521
- Add Qwen3.5 MTP speculative decoding by @aleroot in #351
- Add Qwen 3.5 JSON tool-call fallback by @aleroot in #529
- bump mlx-swift to 0.31.6 -- .4 was missing new API that the current code calls by @davidkoski in #484
- Expose Gemma 4 encoder surface at @_spi(GemmaEncoder) by @xocialize in #530
- Enforce copy semantics on LogitProcessor for speculative decoding (#508) by @aleroot in #533
- Fix tied word embedding weight sanitization by @aleroot in #532
- Add TranslateGemma support (Gemma 3 translation template) by @beshkenadze in #348
- Support video processing options in UserInput and MediaProcessing by @aleroot in #534
- Surface rejected tool calls as generation events by @aleroot in #512
- Add an Xcode 27 compile check and handle Generation.rejectedToolCall by @CharlieTLe in #538
- Fix nested tool grammar schema test by @angelsbrood in #527
- Qwen2.5-VL: accumulate vision cuSeqlens across temporal slices by @NivDvir in #509
- Add Qwen3.5 VLM processor config fallback by @ngutech21 in #546
- Tool call parser hardening by @aleroot in #531
- Muse Glimmer: honor the requested KV-cache capacity by @aleroot in #556
- Report prompt-cache reuse in GenerateCompletionInfo by @aleroot in #559
- Fix Qwen3VL prompt cache reuse for text-only inputs by @ngutech21 in #549
- Add opt-in q4_0 lattice calibration to model conversion by @GoodOlClint in #507
- Add Helium (Kyutai) LLM port by @noorbhatia in #555
- Read Gemma's brace-form object and array values by @aleroot in #557
- Add LoRA dropout and training mode support by @aleroot in #541
- Add reranker API support by @aleroot in #375
- Load weight files a safetensors index leaves out by @aleroot in #562
- Fix MLXFoundationModels compilation against the Xcode 27 beta 5 SDK by @abra-code in #544
- Tune TurboFlash query grouping for short decode by @aleroot in #570
- Fix mixed-precision QLoRA output promotion by @aleroot in #564
- Add extensible VLM processor loading rules by @aleroot in #565
- Share fused Qwen3.5 router top-k across LLM and VLM by @ngutech21 in #567
- Fix SSM gradient transforms across chunks by @aleroot in #571
- MLXVLM: resolve the LFM2VL image token id from the vocabulary by @Siddhesh2377 in #576
- Strengthen chunked SSM gradient tests by @aleroot in #574
- Parallelize weight loading with byte-balanced contiguous groups(up to ~1.8x faster model loads) by @aleroot in #575
- Add variance-normalized KV cache by @aleroot in #329
- Temporarily disable the Xcode 27 integration build on pull requests by @thechriswebb in #582
- Replace MLXRandom.seed with task-local withRandomState in tests by @aleroot in #583
- Add ai usage policy by @ctymoszek in #585
- Generalise compiled decode segments to Qwen 3 Next by @aleroot in #569
- Share fused MoE router top-k across models by @aleroot in #568
- Qwen direct reduction by @aleroot in #573
- Fuse Qwen GDN input projections by @aleroot in #572
- Clear the compiler warnings from the library and both test targets by @thechriswebb in #594
- Leverage GitHub workflows to keep PR labels up-to-date by @thechriswebb in #587
- Qwen3.5: don't shift norm weights just because a checkpoint has MTP tensors by @mandrael in #598
- Fix DeepSeek MTP layer filtering by @aleroot in #592
- Allow downstream specialization of the Qwen3.5 GDN/MoE blocks and SwitchGLU by @agerjura in #511
- Expose Falcon-H1 encoder surface at @_spi(FalconH1Encoder) by @xocialize in #596
- Report prompt token counts to FoundationModels again by @thechriswebb in #591
- Suspend instead of blocking cooperative threads while weights load by @aleroot in #579
- Harden the Foundation Models test fixtures by @thechriswebb in #599
- ParoQuant: extend to MoE and make the whole path fast (#164 follow-up) by @spokvulcan in #471
- Await the async loadWeights overload in the IntegrationTesting tests by @thechriswebb in #605
- Declare module weights as compile state in compiled decode paths by @aleroot in #589
- Gemma4Text: expose decoder layers as loraLayers so MLP LoRA targets load by @cives93 in #602
- Let training differentiate the gated-delta recurrence by @jhancock1975 in #616
- Emit only the new scalars when a token extends the previous character by @spokvulcan in #613
- Gemma 4 VLM: fuse float32 logit softcapping by @aleroot in #615
- Fix generation loop blocking Swift cooperative workers by @aleroot in #611
- Add configuration-based LoRA metadata discovery by @aleroot in #597
- RotatingKVCache: make trim wrap-aware instead of corrupting the ring by @aleroot in #584
- Add bounded cross-dialect tool-call recovery by @aleroot in #548
ChatSession: reuse the KV cache for append-only media turns by @NivDvir in #515- Clear the MLX cache on the first generated token by @negativetime in #620
- Extract the model cache and stop whole-cache eviction abandoning an in-flight load by @thechriswebb in #603
- pick up mlx-swift 0.32.2 by @davidkoski in #646
- Named image attachments in MLXLMCommon and the Foundation Models adapter by @thechriswebb in #535
- Convert the booleans in Foundation Models tool schemas to Swift Bool values by @thechriswebb in #653
- be able to trigger workflows on main. notes about upcoming releases. by @davidkoski in #647
- require mlx-swift-0.32.3 for SDK issue by @davidkoski in #657
- adjustments for running tests on NAX/TF32 hardware by @davidkoski in #649
New Contributors
- @ctymoszek made their first contribution in #385
- @beshkenadze made their first contribution in #346
- @fdagostino made their first contribution in #390
- @Zrxrxrx made their first contribution in #408
- @magicnight made their first contribution in #422
- @gmondada made their first contribution in #321
- @jacobwillemsma made their first contribution in #409
- @sunshinfight made their first contribution in #423
- @SpiraMira made their first contribution in #389
- @CharlieTLe made their first contribution in #427
- @prbme made their first contribution in #405
- @thechriswebb made their first contribution in #438
- @jyauxi made their first contribution in #435
- @xocialize made their first contribution in #387
- @kklimuk made their first contribution in #455
- @codesworth made their first contribution in #479
- @lucasnewman made their first contribution in #322
- @ActuallyTaylor made their first contribution in #519
- @JamesPriceZV made their first contribution in #490
- @ngutech21 made their first contribution in #546
- @noorbhatia made their first contribution in #555
- @abra-code made their first contribution in #544
- @Siddhesh2377 made their first contribution in #576
- @mandrael made their first contribution in #598
- @agerjura made their first contribution in #511
- @cives93 made their first contribution in #602
- @jhancock1975 made their first contribution in #616
- @negativetime made their first contribution in #620
Full Changelog: 3.31.4...3.32.3