Data
Latest data
-
OpenClaw v2026.9.4 Released — Introducing Update Rollback and Integrated Plugin Management
OpenClaw has released v2026.9.4, featuring an automatic rollback function for failed compatible updates, integrated plugin management, and cloud session preparation capabilities.
-
Claude Code v2.1.269 Released — New Plugin Evaluation Features and Enhanced Agent Management in VSCode
Claude Code v2.1.269 has been released. New features include running plugin evaluation suites with reporting, agent map visualization in VSCode, and improved operational stability across various environments.
-
OpenClaw v2026.6.35 — LTS Release Enhancing Security and Reliability
OpenClaw v2026.6.35, the June 2026 Extended Stable (LTS) release, strengthens security boundaries and improves delivery reliability.
-
OpenAI Codex Python SDK v0.154.0 Released — Addition of inference Settings and External Message Support
OpenAI Codex has released Python SDK 0.154.0, introducing new values to control the number of attempts for inference, integration of external messages, and new options for history management.
-
Claude Code v2.1.268 Released — Expanded Administrative Settings and Bug Fixes Across Environments
Anthropic has released Claude Code v2.1.268, introducing integrated cost management via gateway settings, improved functionality for VSCode and Slack integrations, and various bug fixes.
-
NVIDIA Releases DeepSeek-V4-Pro Model quantization in NVFP4 Format
NVIDIA has released a model based on DeepSeek AI's DeepSeek-V4-Pro-0813, quantization in NVFP4 format to improve inference efficiency.
-
vLLM v0.29.0 Released — Model Runner V2 Now Default for All Models, New Model Support and Performance Gains
vLLM has released v0.29.0, which adopts Model Runner V2 as the default for all models and adds support for new models such as Hy4-preview and Qwen3.8-Flash-Next.
-
Codex CLI v0.154.0 Released — Adding GPT-6-Astra Support and Worktree Integration
Codex CLI v0.154.0 has been released, adding GPT-6-Astra to model selection and introducing experimental isolated checkout functionality via worktrees.
-
Claude Code v2.1.267 Released — Effort Level Settings Added and Various Bug Fixes
Claude Code v2.1.267 introduces the maxEffortLevel setting to limit the effort level ceiling for each provider. Additionally, multiple issues related to session resumption, VS Code extensions, and MCP have been resolved.
-
OpenCode v1.18.30 Released — Adds System Prompts for GPT-6 and Fixes Various Providers
OpenCode has released v1.18.30, featuring the addition of Astra system prompts for GPT-6 models and fixes for providers including Bedrock, OpenAI, and Azure.
-
gemini-cli v0.59.0 Released — Security Fixes and Enhanced Workspace Trust
Gemini CLI has released v0.59.0, featuring fixes including the prevention of SSRF during MCP OAuth metadata discovery and the enforcement of workspace trust in restricted mode.
-
mlx-vlm v0.7.0 Released — Expanded Model Support and inference Performance Optimizations
The latest version of MLX-VLM, v0.7.0, has been released. This update adds support for new models such as LLaVA-OneVision and DeepSeek-V4 Flash Vision EXP, alongside multiple optimizations for inference processes.
-
Ollama v0.34.0 Released — Ollama Models Now Available in ChatGPT Desktop
Ollama has released v0.34.0, allowing users to access Ollama models directly from ChatGPT Desktop via the Ollama app for MacOS.
-
openclaw v2026.9.2 Released — Featuring GPT-6 Astra Support and Improved Usability
openclaw has released version 2026.9.2, which includes support for GPT-6 Astra, the ability to change settings without a restart, and enhanced backup functionality.
-
NVIDIA's quantization Model (NVFP4/FP8) of Alibaba's Qwen3.8-27B
NVIDIA is providing a quantization version of the autoregressive language model Qwen3.8-27B developed by Alibaba, using its own Model Optimizer.
-
llama.cpp v0.4.0 Released — New Model Support and Memory Usage Optimizations
llama.cpp has released v0.4.0, featuring initial support for Qwen3.8-Flash-Next and new functionality to prevent memory peaks during loading.
-
codex rust-v0.153.3 Released: GPT-6-Astra Added to Amazon Bedrock and Bug Fixes
codex released rust-v0.153.3 on 2026-09-05. GPT-6-Astra has been added to the model options for Amazon Bedrock, and behavior regarding asynchronous questions has been fixed.
-
claude-code v2.1.261 Released — Fixes Input Bugs and Expands Output Limits
Anthropic has released v2.1.261, the latest version of Claude Code. This update includes fixes for character omission during input, expanded output limits, and improved behavior in proxy environments.
-
openclaw v2026.9.1 Released — Mermaid Diagram Support and Gateway Stability Improvements
openclaw released its latest version, v2026.9.1, on 2026-09-04. The update includes support for Mermaid block rendering, faster installation, and various improvements to Gateway startup and operation.
-
codex rust-v0.153.1 released — Adds support for GPT-6-Astra API configuration
The rust-v0.153.1 release of codex has been made available, adding functionality to configure GPT-6-Astra via API.
-
codex rust-v0.153.0 Released — Vim Mode Undo/Redo Support and Plugin CLI Enhancements
OpenAI has released rust-v0.153.0 of the local LLM development tool "codex." The primary changes include the addition of undo/redo functionality in Vim mode and remote marketplace integration via the plugin CLI.
-
NVIDIA Develops Nemotron-3-Labs-Ultra-Math-SFT Model Specialized in Mathematical inference
NVIDIA has developed Nemotron-3-Labs-Ultra-Math-SFT, a decoder-only Transformer language model specialized in mathematical problem solving and identifying errors in proofs.
-
NVIDIA Releases Nemotron-3-Labs-Ultra-Math-RL Specialized in Mathematical inference
NVIDIA has released Nemotron-3-Labs-Ultra-Math-RL, a decoder-only Transformer model specialized in solving mathematical problems and identifying errors in proofs.
-
NVIDIA Releases FP4quantization Version of Qwen3.8-Flash-Next on Hugging Face
NVIDIA has released a new model repository on Hugging Face featuring Qwen/Qwen3.8-Flash-Next quantization in FP4 (4-bit floating point).
-
Microsoft Research Releases Streaming Speech Recognition Model VibeVoice-ASR-Streaming-7B
Microsoft Research has released VibeVoice-ASR-Streaming, a streaming speech recognition model capable of continuously transcribing who said what in real-time as audio is input.
-
Microsoft Releases Streaming ASR Model Capable of Simultaneous Speaker Diarization and Transcription
Microsoft Research has announced VibeVoice-ASR-Streaming-1.5B, a streaming ASR model that continuously transcribes who said what in real-time as audio is input.
-
ollama v0.33.3 Released — Feature Enhancements Including gemma4 Image and Audio Support via MLX Engine
ollama released v0.33.3 on 2026-09-02. The update includes support for gemma4 image and audio inputs via the MLX engine, along with a new feature to report cached prompt tokens.
-
claude-code v2.1.259 Released — Fixes Parallel Session Conflicts and Bash Permission Control Bugs
Anthropic released v2.1.259 of the local LLM integration tool claude-code on 2026-09-03, fixing major bugs including configuration file corruption in parallel sessions and malfunctions in Bash command Read() denial rules.
-
openclaw v2026.8.2 Released — Linux Desktop Support and Enhanced Upgrade Safety
openclaw has released v2026.8.2, adding desktop applications (.deb/AppImage) for x86-64 Linux and improving configuration retention and session migration reliability during upgrades.
-
llama.cpp b10730 and b10731 Released — Accelerating qwen4exp MTP Inference and Indexer Aggregation
llama.cpp released b10730 and b10731 on September 1, 2026, implementing recurrent state rollback for MTP speculative decoding in qwen4exp models and slicing indexer head aggregation to improve prompt processing speeds.
-
opencode releases v1.18.26 and v1.18.27 in succession, improving Azure sign-in and adjusting timeouts
opencode released v1.18.26 on September 2, 2026, followed by v1.18.27 on the 3rd. These updates include improved stale thinking block resistance for Claude 5, Bedrock bug fixes, changes to the Azure CLI sign-in method, and unified default timeouts for providers and streaming.
-
gemini-cli v0.58.0 Released — 5 Bug Fixes Including Docker Socket Isolation for macOS Seatbelt
gemini-cli released v0.58.0 on 2026-09-02, featuring five bug fixes including Docker socket isolation in macOS Seatbelt environments and corrections to symbolic link evaluation for ignore paths.
-
Anthropic Releases claude-code v2.1.257, Setting Claude Fable 5.1 as Default and Reducing Complex Task Costs by up to 45%
Anthropic has released version 2.1.257 of its developer tool "claude-code," setting the new "Claude Fable 5.1" as the default Fable model. The company explains that the new model improves performance over Fable 5 while being approximately 25% cheaper normally and up to 45% cheaper for complex agentic tasks.
-
codex v0.152.0 Released — Vim Search, MCP Enhancements, and Windows Sandbox Fixes
OpenAI has released version 0.152.0 of codex, its local LLM CLI tool. This update includes the addition of search functionality for Vim mode, MCP-related enhancements, and fixes for the Windows sandbox.
-
DeepSeek Releases V4-Flash-Vision-Exp Model on Hugging Face
On August 31, 2026, DeepSeek released DeepSeek-V4-Flash-Vision-Exp on Hugging Face, a model that generates text from image and text inputs. It is released under the MIT license and supports 8-bit and FP8.
-
llama.cpp b10726–b10729 Released — Performance Improvements for Metal, CUDA, and AVX2 and New Features
On September 1, 2026, four consecutive releases (b10726 to b10729) of llama.cpp were published. These include Metal tuning for M1 Ultra, XOR swizzle flash attention for CUDA, support for quantized concat, and accelerated large-batch processing for IQ models on AVX2.
-
llama-cpp b10724 Released — Significant Speedup in KV Cache Restoration, Improvements to Multiple Backends
KV cache state restoration time has been reduced from 25–63 seconds to 221–424 milliseconds in builds b10718 through b10724. The update also includes WebGPU crash fixes and tuning for ROCm, Metal, and OpenCL.
-
llama-cpp b10712 Released: Vulkan Top-K Acceleration and Improved n-gram Generation Speed
Versions b10707 through b10712 introduce top-k radix select for Vulkan, boosting n-gram generation speed by up to approximately 1.5x. Bug fixes were also applied to hexagon, RPC, and ggml components.
-
mlx-vlm v0.7.0rc0 Released — MTP Support for Qwen3.8-Flash-Next and Numerous Bug Fixes
A release candidate featuring MTP and continuous batching support for Qwen3.8-Flash-Next, an APC redesign, and fixes for tool calling and streaming responses.
-
claude-code v2.1.251 Released — Adds Model Switch Hooks and Security Fixes
Between August 26 and 31, six updates (v2.1.246 to v2.1.252) were released for claude-code. v2.1.251 introduces model switch hooks and live streaming for sub-agents, while fixing a privilege bypass via symbolic links.
-
koboldcpp v1.120 Released: DirectIO Support, New Models, and Bug Fixes
This update adds DirectIO model loading mode, full support for Qwen3.8-Flash-Next and Ling-3.0-flash, and introduces customizable JavaScript tools to Kobold Lite.
-
opencode v1.18.23–25 Released — Multi-provider Bug Fixes and Azure Entra ID Sign-in Support
opencode released three consecutive versions from v1.18.23 to v1.18.25 between August 25 and 28, 2026. Key changes include bug fixes for Azure, Bedrock, and Cloudflare AI Gateway, as well as Microsoft Entra ID sign-in support for the Azure provider.
-
Qwen Releases Qwen-Drive-1.0-4B, a Multimodal Model for Autonomous Driving
On August 27, 2026, the Qwen team released Qwen-Drive-1.0-4B on Hugging Face, a 4B parameter model designed for autonomous driving motion planning and 3D perception.
-
NVIDIA Releases NVFP4 Quantized Model of Muse-Glimmer-30B on Hugging Face
On August 28, 2026, NVIDIA released the NVFP4 quantized build of Muse-Glimmer-30B on Hugging Face. This is an image-and-text input, text-output model released under the Apache 2.0 license.
-
ollama v0.33.2 Released — Dark Mode Restored and MLX Features Added
Two consecutive releases, v0.33.1 and v0.33.2, have been published. Updates include Qwen3.8 Flash Next support and structured output for MLX, along with dark mode restoration and duplicate launch fixes for the macOS app.
-
vLLM v0.28.0 Released: Major Update Featuring Kimi-K3 Optimizations and DeepSeek V4 Support
vLLM v0.28.0 has been released, featuring a major update with 584 commits. Key highlights include performance optimizations for Kimi-K3, sparse MLA support for DeepSeek V4, expanded features in Model Runner V2, and extensions to tiered KV cache offloading.
-
mlx-vlm v0.6.16–17 Released: New Model Support and Metal Resource Leak Fixes
mlx-vlm has released v0.6.16 and v0.6.17. These updates add support for new models such as GLM-5.3-Flash and Qwen3.8-27B, while fixing multiple bugs including Metal buffer leaks and RoPE position errors.
-
NVIDIA Releases FP4 Quantized Model of Qwen3.8-2.4T-A95B on Hugging Face
On August 26, 2026, NVIDIA released the FP4 (4-bit) quantized model "nvidia/Qwen3.8-2.4T-A95B-NVFP4" of Qwen3.8-2.4T-A95B on Hugging Face. Quantized via ModelOpt, it is provided as a model for text generation and conversational use.
-
Z.ai Releases GLM-5.3 Text Generation Model Based on glm_moe_dsa on Hugging Face
On August 25, Z.ai released GLM-5.3, a text generation model utilizing the glm_moe_dsa architecture, on Hugging Face. The model supports English and Chinese and is accompanied by an arXiv paper.
-
Z.ai Releases GLM-5.3-Flash and BF16 Version Models on Hugging Face
On August 25, 2026, Z.ai released two repositories on Hugging Face: the multimodal model GLM-5.3-Flash and its BF16 version.
-
mlx v0.32.2 Released — Floating-Point Bug Fixes and NAX Attention Optimizations
Fixes truncation in floating-point divmod and NaN dropping in median. Adds fused attention paths for NAX devices and memory read optimizations for gqa-8 decoding.
-
gemini-cli v0.59.0-nightly Released — MCP SSRF Prevention and Fail-Closed Trust Control
The v0.59.0-nightly versions of gemini-cli (August 27 to September 1, 2026) include fixes for SSRF prevention in MCP OAuth processing and fail-closed trust control in restricted mode.