MLX-VLM v0.7.2 has been released, bringing expanded support for a wide range of vision-language models and several functional refinements to the library.
MLX-VLM v0.7.2 Release — Expanded model support and enhanced APC management
New Model Support and Improvements
- New Model Support: Added support for Mistral Large 3 (675B VLM), Bonsai 2 27B, LLaDA-Image model family, Meta's Sapiens2 dense prediction models, and Qwen-Image-2.1 (text-to-image and edit).
- Expanded Compatibility: Added support for current moondream2 checkpoints and moondream3 MLX quants, and improved handling for repository variants of supported image models.
- Feature Enhancements: Improved local model discovery and status exposure, and added SAM 3D Objects for MLX.
Bug Fixes and Refinements
- APC (Attention Placement Cache) Management: Refined APC cache layout logic, addressed built-in APC cache layouts without fallback, and improved token counter clarity.
- Model-Specific Fixes: Corrected LFM2-VL/LFM2.5-VL vision tower preprocessing, fixed Qwen3 Omni chunked prefill and combined image/video inputs, and resolved deepseek_v4 weight normalization issues.
- General Fixes: Fixed Anthropic tool call issues regarding empty assistant content and resolved various issues related to moe_offload and audio resampling in load_audio.
Sources
- v0.7.2 (2026-09-21)