Ollama v0.34.1 has been released, bringing performance optimizations and updates to MLX and llama.cpp integration.
Ollama v0.34.1 Released — Performance improvements and MLX updates
New Features and Improvements
- API Performance: The
/api/tagsendpoint is significantly faster when managing large model libraries, with testing showing a reduction in response time from 3.1 seconds to 294 ms in a cold start. - MLX Enhancements: Improved memory handling for MLX on Apple Silicon.
- Model Creation: MLX safetensors are no longer experimental in
ollama create. GGUF model creation now requires using llama.cpp tooling for safetensor conversion and quantization. - Tooling Updates: Updates to both MLX and llama.cpp.
Bug Fixes
- Repeat Token Detection: To reduce false positives (such as those occurring with OCR), runaway repeat token detection now requires 100 repeat tokens to trigger.
- Deprecation: The
typical_pparameter is now deprecated and can no longer be set when creating new models, though support remains for existing GGUF models.
Sources
- v0.34.1 (2026-09-14)