English

Model Releases

vLLM v0.30.0 Release — Expanded Model Support and Model Runner V2 Enhancements

vLLM v0.30.0 has been released, bringing expanded model support, performance enhancements for Model Runner V2, and new features including Fast Start, Watermarking, and HiSparse.

New Features and Improvements

  • Expanded Model Support: Added support for DeepSeek-V4.1-Flash, DeepSeek-V4-Flash-Vision-Exp, GLM-5.3-Flash, K2-Horizon, Cohere Compass, Bailing V3 VL, and Nanbeige4.2 via the Transformers backend.
  • Model Runner V2 Enhancements: Introduced dual-batch overlap in eager mode, full CUDA graphs for microbatched steps, and adaptive verification for speculative decoding.
  • Fast Start: Added a persistent per-GPU weight-cache daemon using CUDA IPC to speed up engine restarts, now supporting FP4 checkpoints and multi-node TP.
  • Watermarking: Implemented Gumbel-max watermarked generation and detection with a keyed PRF and per-request opt-out.
  • HiSparse: Introduced a host-resident tier for sparse-MLA decode to manage GPU memory pressure by spilling KV pages to pinned host memory.
  • Quantization: Added support for targeted online quantization and various new formats including NVFP4 and W4A16.
  • Performance Optimizations: Includes hardware-specific improvements for NVIDIA (Qwen3.8-Flash-Next, Kimi K3), AMD ROCm, Intel XPU, and CPU (DeepSeek-V4).

Bug Fixes

  • Fixed an approximately 2x RL step-time regression related to GPU-compacted sampling masks.
  • Resolved various issues including invalid structured-output requests no longer stopping the engine, and handled absolute bounds for validation-error response bodies to prevent response amplification.

Breaking Changes and Deprecations

  • Scale-out Endpoints: Registration of scale-out endpoints (/render, /derender, /inference/v1/generate) via vllm serve is now opt-in using the --enable-scale-out flag; the VLLM_ENABLE_SCALE_OUT_ENDPOINTS environment variable has been removed.
  • Removed Features: Removed GPTQ activation ordering (g_idx) and deprecated environment variables including VLLM_PREFIX_CACHE_RETENTION_INTERVAL and VLLM_MM_HASHER_ALGORITHM.
  • Deprecated Components: The all Mamba cache mode and the python -m vllm.entrypoints.grpc_server command are now deprecated.
  • YaRN Alignment: Aligned YaRN with Transformers, which may result in different max_model_len values for certain models.

Sources

  1. v0.30.0 (2026-09-22)