Data
Model releases
-
NVIDIA Releases DeepSeek-V4-Pro Model quantization in NVFP4 Format
NVIDIA has released a model based on DeepSeek AI's DeepSeek-V4-Pro-0813, quantization in NVFP4 format to improve inference efficiency.
-
NVIDIA's quantization Model (NVFP4/FP8) of Alibaba's Qwen3.8-27B
NVIDIA is providing a quantization version of the autoregressive language model Qwen3.8-27B developed by Alibaba, using its own Model Optimizer.
-
NVIDIA Develops Nemotron-3-Labs-Ultra-Math-SFT Model Specialized in Mathematical inference
NVIDIA has developed Nemotron-3-Labs-Ultra-Math-SFT, a decoder-only Transformer language model specialized in mathematical problem solving and identifying errors in proofs.
-
NVIDIA Releases Nemotron-3-Labs-Ultra-Math-RL Specialized in Mathematical inference
NVIDIA has released Nemotron-3-Labs-Ultra-Math-RL, a decoder-only Transformer model specialized in solving mathematical problems and identifying errors in proofs.
-
NVIDIA Releases FP4quantization Version of Qwen3.8-Flash-Next on Hugging Face
NVIDIA has released a new model repository on Hugging Face featuring Qwen/Qwen3.8-Flash-Next quantization in FP4 (4-bit floating point).
-
Microsoft Research Releases Streaming Speech Recognition Model VibeVoice-ASR-Streaming-7B
Microsoft Research has released VibeVoice-ASR-Streaming, a streaming speech recognition model capable of continuously transcribing who said what in real-time as audio is input.
-
Microsoft Releases Streaming ASR Model Capable of Simultaneous Speaker Diarization and Transcription
Microsoft Research has announced VibeVoice-ASR-Streaming-1.5B, a streaming ASR model that continuously transcribes who said what in real-time as audio is input.
-
DeepSeek Releases V4-Flash-Vision-Exp Model on Hugging Face
On August 31, 2026, DeepSeek released DeepSeek-V4-Flash-Vision-Exp on Hugging Face, a model that generates text from image and text inputs. It is released under the MIT license and supports 8-bit and FP8.
-
Qwen Releases Qwen-Drive-1.0-4B, a Multimodal Model for Autonomous Driving
On August 27, 2026, the Qwen team released Qwen-Drive-1.0-4B on Hugging Face, a 4B parameter model designed for autonomous driving motion planning and 3D perception.
-
NVIDIA Releases NVFP4 Quantized Model of Muse-Glimmer-30B on Hugging Face
On August 28, 2026, NVIDIA released the NVFP4 quantized build of Muse-Glimmer-30B on Hugging Face. This is an image-and-text input, text-output model released under the Apache 2.0 license.
-
NVIDIA Releases FP4 Quantized Model of Qwen3.8-2.4T-A95B on Hugging Face
On August 26, 2026, NVIDIA released the FP4 (4-bit) quantized model "nvidia/Qwen3.8-2.4T-A95B-NVFP4" of Qwen3.8-2.4T-A95B on Hugging Face. Quantized via ModelOpt, it is provided as a model for text generation and conversational use.
-
Z.ai Releases GLM-5.3 Text Generation Model Based on glm_moe_dsa on Hugging Face
On August 25, Z.ai released GLM-5.3, a text generation model utilizing the glm_moe_dsa architecture, on Hugging Face. The model supports English and Chinese and is accompanied by an arXiv paper.
-
Z.ai Releases GLM-5.3-Flash and BF16 Version Models on Hugging Face
On August 25, 2026, Z.ai released two repositories on Hugging Face: the multimodal model GLM-5.3-Flash and its BF16 version.