Alibaba has released the weights for the multimodal MoE model "Qwen3.8-Flash-Next" on Hugging Face and ModelScope. This model serves to provide the community with an early preview of the architecture intended for use in the Qwen4 series, similar to the role Qwen3-Next played relative to Qwen3.5.

On the technical side, improvements have been made in four areas: attention, residuals, embeddings, and optimization. The main model consists of 125B parameters, with an additional 51B parameters for N-gram embeddings. It is designed so that 6B parameters are activated per token.

Compared to Qwen3.7-Plus, training costs have been reduced to approximately 1/9. Capabilities in coding and office tasks have improved. It natively supports a context window of 262,144 tokens, which can be extended up to 1 million tokens via YaRN.

The production version will be available on QwenCloud as "Qwen3.8-Flash." It features a 1M context window by default and includes integrated official tools. Pricing is set at $0.15 per million input tokens and $0.47 per million output tokens.


Source: Qwen3.8-Flash-Next (HN 701pt, 233 comments) (HN Search (backfill), 2026-08-26)