Data

Inference runtimes

  1. DATA / Inference runtimes

    OpenClaw v2026.9.4 Released — Introducing Update Rollback and Integrated Plugin Management

    OpenClaw has released v2026.9.4, featuring an automatic rollback function for failed compatible updates, integrated plugin management, and cloud session preparation capabilities.

  2. DATA / Inference runtimes

    OpenClaw v2026.6.35 — LTS Release Enhancing Security and Reliability

    OpenClaw v2026.6.35, the June 2026 Extended Stable (LTS) release, strengthens security boundaries and improves delivery reliability.

  3. DATA / Inference runtimes

    vLLM v0.29.0 Released — Model Runner V2 Now Default for All Models, New Model Support and Performance Gains

    vLLM has released v0.29.0, which adopts Model Runner V2 as the default for all models and adds support for new models such as Hy4-preview and Qwen3.8-Flash-Next.

  4. DATA / Inference runtimes

    Ollama v0.34.0 Released — Ollama Models Now Available in ChatGPT Desktop

    Ollama has released v0.34.0, allowing users to access Ollama models directly from ChatGPT Desktop via the Ollama app for MacOS.

  5. DATA / Inference runtimes

    openclaw v2026.9.2 Released — Featuring GPT-6 Astra Support and Improved Usability

    openclaw has released version 2026.9.2, which includes support for GPT-6 Astra, the ability to change settings without a restart, and enhanced backup functionality.

  6. DATA / Inference runtimes

    llama.cpp v0.4.0 Released — New Model Support and Memory Usage Optimizations

    llama.cpp has released v0.4.0, featuring initial support for Qwen3.8-Flash-Next and new functionality to prevent memory peaks during loading.

  7. DATA / Inference runtimes

    openclaw v2026.9.1 Released — Mermaid Diagram Support and Gateway Stability Improvements

    openclaw released its latest version, v2026.9.1, on 2026-09-04. The update includes support for Mermaid block rendering, faster installation, and various improvements to Gateway startup and operation.

  8. DATA / Inference runtimes

    ollama v0.33.3 Released — Feature Enhancements Including gemma4 Image and Audio Support via MLX Engine

    ollama released v0.33.3 on 2026-09-02. The update includes support for gemma4 image and audio inputs via the MLX engine, along with a new feature to report cached prompt tokens.

  9. DATA / Inference runtimes

    openclaw v2026.8.2 Released — Linux Desktop Support and Enhanced Upgrade Safety

    openclaw has released v2026.8.2, adding desktop applications (.deb/AppImage) for x86-64 Linux and improving configuration retention and session migration reliability during upgrades.

  10. DATA / Inference runtimes

    llama.cpp b10730 and b10731 Released — Accelerating qwen4exp MTP Inference and Indexer Aggregation

    llama.cpp released b10730 and b10731 on September 1, 2026, implementing recurrent state rollback for MTP speculative decoding in qwen4exp models and slicing indexer head aggregation to improve prompt processing speeds.

  11. DATA / Inference runtimes

    llama.cpp b10726–b10729 Released — Performance Improvements for Metal, CUDA, and AVX2 and New Features

    On September 1, 2026, four consecutive releases (b10726 to b10729) of llama.cpp were published. These include Metal tuning for M1 Ultra, XOR swizzle flash attention for CUDA, support for quantized concat, and accelerated large-batch processing for IQ models on AVX2.

  12. DATA / Inference runtimes

    llama-cpp b10724 Released — Significant Speedup in KV Cache Restoration, Improvements to Multiple Backends

    KV cache state restoration time has been reduced from 25–63 seconds to 221–424 milliseconds in builds b10718 through b10724. The update also includes WebGPU crash fixes and tuning for ROCm, Metal, and OpenCL.

  13. DATA / Inference runtimes

    llama-cpp b10712 Released: Vulkan Top-K Acceleration and Improved n-gram Generation Speed

    Versions b10707 through b10712 introduce top-k radix select for Vulkan, boosting n-gram generation speed by up to approximately 1.5x. Bug fixes were also applied to hexagon, RPC, and ggml components.

  14. DATA / Inference runtimes

    koboldcpp v1.120 Released: DirectIO Support, New Models, and Bug Fixes

    This update adds DirectIO model loading mode, full support for Qwen3.8-Flash-Next and Ling-3.0-flash, and introduces customizable JavaScript tools to Kobold Lite.

  15. DATA / Inference runtimes

    ollama v0.33.2 Released — Dark Mode Restored and MLX Features Added

    Two consecutive releases, v0.33.1 and v0.33.2, have been published. Updates include Qwen3.8 Flash Next support and structured output for MLX, along with dark mode restoration and duplicate launch fixes for the macOS app.

  16. DATA / Inference runtimes

    vLLM v0.28.0 Released: Major Update Featuring Kimi-K3 Optimizations and DeepSeek V4 Support

    vLLM v0.28.0 has been released, featuring a major update with 584 commits. Key highlights include performance optimizations for Kimi-K3, sparse MLA support for DeepSeek V4, expanded features in Model Runner V2, and extensions to tiered KV cache offloading.