On August 26, 2026, NVIDIA released the model repository "nvidia/Qwen3.8-2.4T-A95B-NVFP4" on Hugging Face. The base model is Qwen/Qwen3.8-2.4T-A95B, and it is provided as a quantized build using FP4 (4-bit precision) via NVIDIA's ModelOpt (Model Optimizer).

According to the tag information, the model is classified as a text generation model with a MoE architecture (qwen3_5_moe_text), and "conversational" is specified as its intended use. The weights are stored in safetensors format, and the license is categorized as "other."


Source: nvidia/Qwen3.8-2.4T-A95B-NVFP4 (HF: NVIDIA, 2026-08-26)