English

Pricing & LimitsXiaomiMiMo-v2.5

Xiaomi Slashes MiMo-v2.5 Series API Prices by Up to 99%, Releases 1 Trillion Parameter Model with 1M Context

This article is a translation. Read the Japanese original

On the 27th, Xiaomi announced that it will reduce API pricing for the MiMo-v2.5 series by up to 99%. This reduction applies to all models in the series, with monthly and annual subscription plans available.

The model architecture features 1 trillion total parameters, 42B active parameters, and a 1M context window. The company claims that the models demonstrate performance comparable to Claude Opus 4.6 in high-intensity agent workloads. To improve inference throughput, Xiaomi explained that it combined FP4 lossless quantization with DFlash parallel speculative decoding, and for GPU efficiency, it adopted persistent kernels and heterogeneous pipeline optimization via TileRT.

In terms of functionality, the series features omni-modal understanding that natively processes images, video, audio, and text, as well as long-context reasoning with 1M context and agent execution (browsing, reasoning, and operation). For speech synthesis, it supports fine-grained control over speaking speed, emotion, and tone, as well as the generation of new timbres from a single sentence and cloning from a few samples. For speech recognition, the company claims robustness in high-noise, far-field, and multi-speaker environments, as well as support for Chinese-English bilingualism, Chinese dialects, and lyric transcription.

For developers, OpenAI and Anthropic protocol-compatible APIs and low-latency inference services are provided. For coding agent use cases, it achieves a peak speed of 1,000 tokens per second, and the free tier has been expanded from 1 hour to 4 hours per day. The flagship mimo-v2.5-pro features native integration with OpenClaw and collaboration with the Kingsoft Office (WPS) ecosystem.


Source: Xiaomi MiMo-v2.5 Series API Permanent Price Reduction Up to 99%(HN 134pt・160コメント) (HN Search (backfill), 2026-05-27)