What's Changed
- Fix scan and sort for a zero-size axis by @eyupcanakman in #4340
- Add a correction parameter in std and var by @prady0t in #4348
- Fix tests/run.py stuck in macOS CI by @zcbenz in #4410
- Extension Example Updated by @jagrit06 in #4412
- Fix sorted gather_qmm NAX row overflow above 32K by @PhilipJohnBasile in #3922
- Fix deadlock caused by mx.clear_streams() holding GIL by @aleroot in #4413
- Fix integer pow zeroing a whole SIMD vector on a negative exponent by @ayaangazali in #4354
- Make concurrency cap on load adaptative so that I/O scales with the machine by @aleroot in #4408
- Fix pad with an axes subset and negative axes by @kapellirohith in #4364
- Use sdpa_vector_2pass_1_gqa for GQA size 12 and 16 by @dudududukim in #4380
- python: Make axis of put_along_axis optional by @ayaangazali in #4360
- python: Make axis default to -1 in take_along_axis by @aaishwarymishra in #4368
- Propagate NaN in arg reductions by @atirna in #4291
- Fix mx.from_fp8 E4M3FN NaN decode with branchless carry by @saud5150 in #4376
- Update rules for PR limitation bypass list by @zcbenz in #4415
- [CUDA] Use native events for GPU fence waits by @strayberry in #4401
- Validate GGUF tensor dimensions by @roshaninfordham in #4378
- Fix crash on ellipsis indexing with too many trailing indices by @Adityaj0 in #4396
- [CUDA] Cholesky via cuSOLVER by @sashko-zakharchuk in #4208
- python: Fix setitem with negative index preceded by None by @Adityaj0 in #4414
- Update nanobind to 3.0.1 by @XXXXRT666 in #4417
- chore: Use normalize_axis_index in scan ops by @Adityaj0 in #4383
- python: Add explicit index and bytes support by @aaishwarymishra in #4388
- docs: fix typo recived -> received by @vaibhav8a in #4424
- Fix Metal grid sizing for strided scans by @TheDarkchip in #4420
- Fix SGD and Adafactor weight decay mutating caller arrays by @2sumtech in #4426
- Widen float16/bfloat16 to float32 in cpu reduce by @Ved235 in #4387
- sdpa_vector_2pass_1_gqa kernel batch offset for K/V in gqa decode by @dudududukim in #4431
- Make head-dim-256 prefill with array mask run fused NAX kernel by @dwijenpatel in #4416
- python: Make unstack return tuple instead of list by @JasonHonKL in #4448
- python: Make iter() throw TypeError for 0-dim array by @simeetnayan81 in #4425
- Fix save and save_safetensors corrupting a lazily loaded source file by @Cdimoy in #4434
- [Metal] Use hypot for complex ops by @JasonHonKL in #4345
- Add get_array_buffer_size to query buffer size for arrays by @aleroot in #4436
- Add M5 ultra tunings for non-quantized matmuls by @jagrit06 in #4447
- [CUDA] Fix completion worker busy loop by @strayberry in #4452
- Fix non transposed affine qmm dispatch logic by @RohanGautam in #4392
- Fix sorted gather_qmm on ragged K by @erwinzhang7 in #4009
- Don't restore a thread-affine stream when StreamContext dies on another thread by @michalk8 in #4462
- Fix Pad vjp with an axes subset and negative axes by @nileshpatil6 in #4441
- Fix integer div-by-zero crash in CPU backend (sort, scan, binary ops) by @MarcosAsh in #4442
- Use 64-bit file seeks on Windows by @dhiltgen in #4456
- Use precise::exp in Sigmoid so compiled and eager sigmoid agree by @pierre427 in #4461
- Use NAX attention for long unmasked D72/D80 inputs by @wyanzhao in #4455
- Add D512 support to Metal vector attention by @wyanzhao in #4459
- Add a scoped_env utility for tests by @zcbenz in #4474
- [Metal] global scale for qmm by @nastya236 in #4458
- Fix re-entering the same stream context manager by @michalk8 in #4478
- leak global CommandEncoder -- avoid cuda synchronize on process shutdown by @davidkoski in #4480
- Update CONTRIBUTING.md by @zcbenz in #4469
- python: fix bytes() on non-contiguous arrays by @axiom-of-choice in #4449
- Break the sibling cycle when an array is released by assignment by @tudalex in #4453
- Win: Use DXGI for accurate WDDM VRAM memory budget by @dhiltgen in #4457
- Add type hints and docstrings for ALiBi layer by @Ritabanm in #4472
- Fix fp quantized matmul corruption when the quantized dim is not a multiple of 32 by @kapellirohith in #3912
- Fix CUDA test synchronization flake by @dhiltgen in #4490
- Skip bypass list updates in forks by @XXXXRT666 in #4495
- Load global scales in qmm_t kernels by @dhiltgen in #4483
- Release the GIL when copying Metal DLPack inputs (mx - pytorch mps bridge) by @WindChimeRan in #4497
- Use matrix kernels for global-scale gather_qqmm by @dhiltgen in #4481
- Adding metal kernels for the gated delta nets. by @tpegolotti in #4020
- Fix jaccl ring all_gather: direction 1 slice was not mirrored by @Drifter4242 in #4443
- Skip Metal-only gated delta kernel tests on other backends(fix CI) by @aleroot in #4522
- [CUDA] Add global scale support to gather_qmm by @dhiltgen in #4507
- [BUG][Metal] Deadlock in fence by @nastya236 in #4552
- Fix ordering of complex64 with NaNs by @louen in #4519
- Improve metal memory usage for SDPA D256 by @dhiltgen in #4505
- Add support for matrix transpose by @aaishwarymishra in #4402
- Improve metal memory usage for SDPA D512 by @dhiltgen in #4487
- Support shapeless compilation of scan operations by @keeeeenw in #4510
- [Metal] Gather mm improvement by @nastya236 in #4567
- restoring perf regression for upsample
linearmode align_corners=False by @sp4s-s in #4500 - Add fused Metal kernels for fast.cross_entropy by @zsun6 in #4520
- Add missing defaults to tri, tril, triu, gather_mm signatures by @rishabhsai in #4492
- Add Metal SDPA support for D96/V64 by @wyanzhao in #4499
- Raise IndexError for out of bounds axes by @devangpratap in #4484
- Fix floor_divide for integers by @aaishwarymishra in #4515
- Fix cpu binary ops on large data by @Prudctual in #4517
- Do not throw from CUDA destructors and avoid implicit default streams in eval/compile by @aleroot in #4514
- [Metal] gather_qmm improvement by @nastya236 in #4572
- Fix floating-point constant precision in compiled kernels by @Ryan11c in #4511
- chore: Fix assertion in ReduceScatter::eval_cpu by @ronaldmannak in #4557
- Leak thread local streams on exit in main thread by @zcbenz in #4576
New Contributors
- @atirna made their first contribution in #4291
- @saud5150 made their first contribution in #4376
- @roshaninfordham made their first contribution in #4378
- @vaibhav8a made their first contribution in #4424
- @TheDarkchip made their first contribution in #4420
- @2sumtech made their first contribution in #4426
- @simeetnayan81 made their first contribution in #4425
- @Cdimoy made their first contribution in #4434
- @michalk8 made their first contribution in #4462
- @MarcosAsh made their first contribution in #4442
- @tudalex made their first contribution in #4453
- @Ritabanm made their first contribution in #4472
- @tpegolotti made their first contribution in #4020
- @Drifter4242 made their first contribution in #4443
- @keeeeenw made their first contribution in #4510
- @sp4s-s made their first contribution in #4500
- @zsun6 made their first contribution in #4520
- @rishabhsai made their first contribution in #4492
- @devangpratap made their first contribution in #4484
- @Prudctual made their first contribution in #4517
- @Ryan11c made their first contribution in #4511
- @ronaldmannak made their first contribution in #4557
Full Changelog: v0.32.2...v0.32.3