MLX-VLM v0.7.4 has been released, bringing support for several new models and significant improvements to image generation workflows.
Model ReleasesDeepSeek V4.1Qwen3.5 MoEInternVL3.5
MLX-VLM v0.7.4 Public — Adds support for DeepSeek V4.1 and multiple image generation enhancements
New Features and Improvements
- Added support for DeepSeek V4.1 (vision and language models) and InternVL3.5.
- Introduced the ability to generate multiple images per prompt via the
num_imagesparameter. - Implemented batched image generation for
qwen_image,flux2, andideogram4. - Added support for Qwen3.5 MoE text and LFM2 encoder/ColBERT.
- Added support for
gpt-osswith reasoning and content parsing for harmony channels. - Improved image generation by using model-based sampling defaults.
Bug Fixes
- Fixed audio request failures for MiMo-V2.6 in server multi-threaded environments.
- Fixed Qwen3.5-arch left padding issues for batches and Qwen3-VL position embeddings with quantized
pos_embed. - Resolved tool-parser failures involving boolean property schemas, untyped Qwen3-Coder objects, and unclosed parameters.
- Fixed issues with sampling for small
top-pand very low temperature values. - Fixed a bug in the Gemma 4 audio encoder regarding past frame attention.
- Resolved issues with interleaved image order in shared prompt formatting and KV cache packing for certain quantized models.
Sources
- v0.7.4 (2026-09-28)