English

Product LaunchesOllama

Ollama v0.34.1 Released — Performance improvements and MLX updates

Ollama v0.34.1 has been released, bringing performance optimizations and updates to MLX and llama.cpp integration.

New Features and Improvements

  • API Performance: The /api/tags endpoint is significantly faster when managing large model libraries, with testing showing a reduction in response time from 3.1 seconds to 294 ms in a cold start.
  • MLX Enhancements: Improved memory handling for MLX on Apple Silicon.
  • Model Creation: MLX safetensors are no longer experimental in ollama create. GGUF model creation now requires using llama.cpp tooling for safetensor conversion and quantization.
  • Tooling Updates: Updates to both MLX and llama.cpp.

Bug Fixes

  • Repeat Token Detection: To reduce false positives (such as those occurring with OCR), runaway repeat token detection now requires 100 repeat tokens to trigger.
  • Deprecation: The typical_p parameter is now deprecated and can no longer be set when creating new models, though support remains for existing GGUF models.

Sources

  1. v0.34.1 (2026-09-14)