English

Releases

MLX v0.32.3 Release — Extensive bug fixes and backend performance improvements

MLX has released version 0.32.3, which focuses on enhancing stability and performance across its various execution backends, including CPU, Metal, and CUDA.

Key Improvements and New Features

  • Metal Optimizations: Added support for D512 in Metal vector attention, introduced Metal kernels for gated delta nets, and implemented fused Metal kernels for fast.cross_entropy.
  • Performance and Scaling: Made concurrency caps for loading adaptive to allow I/O to scale with the machine and added M5 Ultra tunings for non-quantized matmuls.
  • CUDA Enhancements: Implemented native events for GPU fence waits and added global scale support to gather_qmm.
  • Python API Improvements: Made the axis parameter optional in put_along_axis and updated unstack to return tuples instead of lists.

Bug Fixes

  • Backend Stability: Fixed several critical issues, including a deadlock caused by mx.clear_streams() holding the GIL, integer division-by-zero crashes in the CPU backend, and memory usage improvements for SDPA (Scaled Dot-Product Attention) in Metal.
  • Data Integrity: Addressed issues regarding NaN propagation in arg reductions, floating-point constant precision in compiled kernels, and corruption when saving lazily loaded files.
  • General Fixes: Resolved issues with GGUF tensor dimension validation, ellipsis indexing crashes, and several macOS CI-related hangs.

Sources

  1. v0.32.3 (2026-09-29)