mlx-vlm v0.7.1 has been released, bringing expanded support for multiple vision-language models and several architectural optimizations.
mlx-vlm v0.7.1 Public — Expanded model support and performance optimizations
New Features and Improvements
- Expanded Model Support: Added support for GLM-5 Next, Hy4, Spark-X2.5, and PP-DocLayout V3 (MLX).
- Vision & Video Capabilities: Added standalone DINOv2 model with register tokens and mask support, SAM 3.1 multiplex video tracking port, and video understanding support for Mage-VL.
- Performance Optimizations: Optimized Gemma 4 KV-shared prefill and accelerated Qwen MXFP4 verification.
- Other Features: Added TTS support for MiniCPM-o and Indic-OCR support.
Bug Fixes
- Fixed MiniCPM5 function-call parsing and Mllama cross-attention masks during chunked prefill.
- Resolved issues with Gemma 4 moe_offload nested Experts.switch_glu paths.
- Fixed issues related to Qwen model embedding, and resolved marker stripping for GLM and Aya models.
Sources
- v0.7.1 (2026-09-14)