Ollama has released version 0.34.4, which brings performance optimizations for specific models and stability improvements across platforms.
Ollama v0.34.4 Release — Faster thinking models and improved Apple Silicon support
Key improvements and new features include:
- Structured outputs for thinking models are now applied in a single pass, making them faster and more reliable.
- Qwen 3.8 prompt processing is faster on Apple Silicon.
- Gemma 4 on Apple Silicon now picks the best image resolution per image, keeping more detail in high-resolution images.
- Updated llama.cpp, MLX, and XGrammar.
The update also addresses the following issues:
- Fixed intermittent "model not found" errors with a large local library.
- Fixed the macOS app becoming unresponsive when checking if ChatGPT or Codex is running.
Sources
- v0.34.4 (2026-09-23)