MLX has released version 0.32.3, which focuses on enhancing stability and performance across its various execution backends, including CPU, Metal, and CUDA.
MLX v0.32.3 Release — Extensive bug fixes and backend performance improvements
Key Improvements and New Features
- Metal Optimizations: Added support for D512 in Metal vector attention, introduced Metal kernels for gated delta nets, and implemented fused Metal kernels for
fast.cross_entropy. - Performance and Scaling: Made concurrency caps for loading adaptive to allow I/O to scale with the machine and added M5 Ultra tunings for non-quantized matmuls.
- CUDA Enhancements: Implemented native events for GPU fence waits and added global scale support to
gather_qmm. - Python API Improvements: Made the
axisparameter optional input_along_axisand updatedunstackto return tuples instead of lists.
Bug Fixes
- Backend Stability: Fixed several critical issues, including a deadlock caused by
mx.clear_streams()holding the GIL, integer division-by-zero crashes in the CPU backend, and memory usage improvements for SDPA (Scaled Dot-Product Attention) in Metal. - Data Integrity: Addressed issues regarding NaN propagation in arg reductions, floating-point constant precision in compiled kernels, and corruption when saving lazily loaded files.
- General Fixes: Resolved issues with GGUF tensor dimension validation, ellipsis indexing crashes, and several macOS CI-related hangs.
Sources
- v0.32.3 (2026-09-29)