vLLM v0.30.0 has been released, bringing expanded model support, performance enhancements for Model Runner V2, and new features including Fast Start, Watermarking, and HiSparse.
vLLM v0.30.0 Release — Expanded Model Support and Model Runner V2 Enhancements
New Features and Improvements
- Expanded Model Support: Added support for DeepSeek-V4.1-Flash, DeepSeek-V4-Flash-Vision-Exp, GLM-5.3-Flash, K2-Horizon, Cohere Compass, Bailing V3 VL, and Nanbeige4.2 via the Transformers backend.
- Model Runner V2 Enhancements: Introduced dual-batch overlap in eager mode, full CUDA graphs for microbatched steps, and adaptive verification for speculative decoding.
- Fast Start: Added a persistent per-GPU weight-cache daemon using CUDA IPC to speed up engine restarts, now supporting FP4 checkpoints and multi-node TP.
- Watermarking: Implemented Gumbel-max watermarked generation and detection with a keyed PRF and per-request opt-out.
- HiSparse: Introduced a host-resident tier for sparse-MLA decode to manage GPU memory pressure by spilling KV pages to pinned host memory.
- Quantization: Added support for targeted online quantization and various new formats including NVFP4 and W4A16.
- Performance Optimizations: Includes hardware-specific improvements for NVIDIA (Qwen3.8-Flash-Next, Kimi K3), AMD ROCm, Intel XPU, and CPU (DeepSeek-V4).
Bug Fixes
- Fixed an approximately 2x RL step-time regression related to GPU-compacted sampling masks.
- Resolved various issues including invalid structured-output requests no longer stopping the engine, and handled absolute bounds for validation-error response bodies to prevent response amplification.
Breaking Changes and Deprecations
- Scale-out Endpoints: Registration of scale-out endpoints (
/render,/derender,/inference/v1/generate) viavllm serveis now opt-in using the--enable-scale-outflag; theVLLM_ENABLE_SCALE_OUT_ENDPOINTSenvironment variable has been removed. - Removed Features: Removed GPTQ activation ordering (
g_idx) and deprecated environment variables includingVLLM_PREFIX_CACHE_RETENTION_INTERVALandVLLM_MM_HASHER_ALGORITHM. - Deprecated Components: The
allMamba cache mode and thepython -m vllm.entrypoints.grpc_servercommand are now deprecated. - YaRN Alignment: Aligned YaRN with Transformers, which may result in different
max_model_lenvalues for certain models.
Sources
- v0.30.0 (2026-09-22)