English

Model ReleasesDeepSeek V4.1Qwen3.5 MoEInternVL3.5

MLX-VLM v0.7.4 Public — Adds support for DeepSeek V4.1 and multiple image generation enhancements

MLX-VLM v0.7.4 has been released, bringing support for several new models and significant improvements to image generation workflows.

New Features and Improvements

  • Added support for DeepSeek V4.1 (vision and language models) and InternVL3.5.
  • Introduced the ability to generate multiple images per prompt via the num_images parameter.
  • Implemented batched image generation for qwen_image, flux2, and ideogram4.
  • Added support for Qwen3.5 MoE text and LFM2 encoder/ColBERT.
  • Added support for gpt-oss with reasoning and content parsing for harmony channels.
  • Improved image generation by using model-based sampling defaults.

Bug Fixes

  • Fixed audio request failures for MiMo-V2.6 in server multi-threaded environments.
  • Fixed Qwen3.5-arch left padding issues for batches and Qwen3-VL position embeddings with quantized pos_embed.
  • Resolved tool-parser failures involving boolean property schemas, untyped Qwen3-Coder objects, and unclosed parameters.
  • Fixed issues with sampling for small top-p and very low temperature values.
  • Fixed a bug in the Gemma 4 audio encoder regarding past frame attention.
  • Resolved issues with interleaved image order in shared prompt formatting and KV cache packing for certain quantized models.

Sources

  1. v0.7.4 (2026-09-28)