Xiaomi announced on May 27, 2026, that it has launched the Token Plan for its AI model service "MiMo" globally. The plan covers V2.5 series models through monthly and annual subscriptions, and the free usage quota has been expanded from 1 to 4 hours per day. The company states that peak inference speeds reach 1,000 tokens/s.
Product LaunchesPricing & LimitsXiaomiMiMomimo-v2.5-pro
Xiaomi Launches MiMo Token Plan Globally; mimo-v2.5-pro Features 1T Total Parameters and 1M Context
This article is a translation. Read the Japanese original
The flagship mimo-v2.5-pro features a configuration of 1T total parameters, 42B active parameters, and a 1M context window, with enhanced OpenClaw integration and support for the WPS ecosystem. The company claims that it demonstrates performance comparable to Claude Opus 4.6 in high-load agent-based workloads. For inference, it employs FP4 lossless quantization and DFlash parallel speculative decoding, while utilizing TileRT persistent kernels and heterogeneous pipeline optimization for GPU efficiency.
As an omni-modal understanding system, it natively processes images, video, audio, and text, and includes agent capabilities to autonomously execute browsing, reasoning, and operations. Speech synthesis supports fine-grained control of speaking rate, emotion, and tone, as well as new voice generation from a single sentence and cloning with small samples. Speech recognition handles Chinese-English bilingualism, various Chinese dialects, and lyric transcription, and is said to operate stably even in high-noise, far-field, and multi-speaker environments.
For developers, the company provides a platform featuring OpenAI and Anthropic compatible APIs, low-latency inference, and comprehensive development documentation.
Source: Xiaomi MiMo Token Plan is Now Globally Available(HN 79pt・57コメント)(HN Search (backfill), 2026-05-27)