llama.cpp has released version 0.5.0, focusing on backend performance, correctness, and broader model coverage. The update includes optimizations for CUDA and Metal, more robust server operations, and support for several new model architectures.
llama.cpp v0.5.0 Public Release — Backend performance optimizations and expanded model support
New Features and Improvements
- Performance Optimizations:
- Accelerated CUDA
conv2dusing implicit GEMM. - Added Metal MoE and SSM_CONV fusion optimizations.
- Enabled CUDA graphs for MTP drafting.
- Accelerated CUDA
- New Model Support:
- Added support for HRM-Text (DFM Mimir 1B).
- Added conversion support for MiMo-V2.6 and support for HunyuanOCR via DFlash.
- Extended support for Nemotron MTP and Nemotron-H models.
- Added Qwen4Exp hyper-connection operations and sparse flash attention.
- Server and API Enhancements:
- Allowed the server to bind to multiple addresses via
--host. - Added
input_imagesupport to server function-call outputs. - Improved JSON Schema and PEG handling.
- Added environment variables for temperature, top-p, min-p, and penalties.
- Allowed the server to bind to multiple addresses via
Bug Fixes
- Fixed token counting API crashes during sleep in the server.
- Fixed tensor-parallel split state and granularity issues for fused QKV models.
- Fixed SigLIP buffer overruns for tall/wide images.
- Resolved various UI issues, including mobile breakpoint and content overflow problems.
- Fixed several parser errors for Ling 3.0, DeepSeek V3.2/V4, qwen3-coder, Muse Glimmer, and Gemma 4.
Sources
- v0.5.0 (2026-09-23)