0.32.2
IMPORTANT
Use release 0.32.3 -- important patch for #491
Highlights
- New mlx release, https://github.com/ml-explore/mlx/compare/v0.31.1...v0.32.2
- Distributed support by @ronaldmannak in #482
- replaced and expanded integration tests - #477
- add thin logging API - #484
API Changes Changes
Device.deviceTypeis no longer optionalDeviceequality now compares indexDevice.gpuandDevice.cpunow return the current device per theStream, e.g. on a multi-gpu CUDA system it might return something other than device 0Stream.gpuandStream.cpunow return the current GPU or CPU stream (task aware)Stream.defaultStream(new) returns the current streamMLXNN.Lion.betasdefault values changed (now correct -- matching python)
Deprecations
StreamOrDevice.device(_:)→ prefer.stream(Stream)Stream.init()andStream.init(_ device:)→ useStream.withNewDefaultStreamStream.defaultStream(_ device: Device)→ usedefaultStream(DeviceType)/withNewDefaultStream(Device)Device.setDefaultDevicemessage tightened: only valid before any mlx resources are created
Behavior Changes
tensordot(_:_:axes:)default axes: 1 → 2 (matches Python).nanToNum(posInf:negInf:)defaults 0 → nil (now replaces with dtype max/min, matching Python).linspacenow takes dtype:; integer args default to float32 (previously T.dtype), Double args resolve to float32 instead of float64.- convolve
.samemode: padRight fix for even kernels (padLeft/2 - 1 → padLeft - 1). MLXNN.Poolpadding widths fix (was emitting an extra leading/trailing pair).ALiBi.alibiSloperewritten to match Python for non-power-of-2 head counts; alibiMatrix x2 axis fix (real distance matrix now).GRU: hidden bias bhn now applied when there is no incoming hidden state.Adafactor: uses broadcast instead of matmul (works for rank > 2), and lazily creates state instead of force-unwrapping.Lion: default beta2 0.999 → 0.99.Muon: Newton–Schulz now uses norm/addMM — small numeric drift vs. before.Device/Streamsemantics: streams are now pooled and default device/stream are task-local pairs; Device.cpu/.gpu reflect the enclosing withDefaultDevice/withNewDefaultStream scope rather than fixed index-0 devices.
What's Changed
- Add fp8 conversion utilities by @lucasnewman in #435
- build iOS in CI by @davidkoski in #440
- Add global scale support to quantized layers by @aleroot in #426
- Fix FinalizerCaptureState leak in MLXArray(rawPointer:_:dtype:finalizer:) by @agerjura in #448
- Derive the jit-source lists in update-mlx.sh instead of hand-listing them by @GoodOlClint in #445
- Improve CUDA NVCC spawn performance (on Linux) by @Joannis in #451
- Condition CUDA build plugin on Linux hosts by @ThorFuchs in #447
- Nested initializers by @louen in #456
- [CUDA] Build the CUDA argument lists in steps by @GoodOlClint in #442
- Fix a deadlock between
CompiledFunction.lockandevalLock, and a data race on the compiler cache by @DanielMercado60660 in #461 - Fix StreamOrDevice.stream(_:) ignoring its argument by @aleroot in #463
- test Module parameters and compile by @davidkoski in #465
- update for mlx v0.32.2 by @CharlieTLe in #450
- Fix wired-memory ticket lifecycle during cancellation by @aleroot in #471
- Fix typo condiiton -> condition by @Silenterc in #473
- fix(ops): fix convolve padRight for even kernels in same mode by @shubhransh-gupta in #467
- Add countNonzero to the Swift API by @Silenterc in #479
- Distributed support by @ronaldmannak in #482
- add thin logging shim by @davidkoski in #484
- patch for Cuda builds by @davidkoski in #478
- pool Streams, fix Device inheritance by @davidkoski in #472
- replace and improve integration tests by @davidkoski in #477
- pick up fix-custom-io from mlx-c by @davidkoski in #476
- package: exclude mlx/build cmake artifacts from Cmlx target by @Nicolas-nwb in #402
New Contributors
- @agerjura made their first contribution in #448
- @ThorFuchs made their first contribution in #447
- @DanielMercado60660 made their first contribution in #461
- @Silenterc made their first contribution in #473
- @shubhransh-gupta made their first contribution in #467
Full Changelog: 0.31.6...0.32.2