Xiaomi has released Xiaomi-Robotics-1, a robot policy model that combines embodiment-free (UMI) pre-training with post-training using real robot data. Adopting a two-stage learning process similar to LLMs, the model is designed to acquire general representations for action generation during pre-training, followed by post-training to align the model with real robots and instruction following.
Model ReleasesXiaomiXiaomi-Robotics-1
Xiaomi Releases Xiaomi-Robotics-1 Robot Policy Model Trained with 100,000 Hours of Embodiment-Free Pre-training
This article is a translation. Read the Japanese original
For pre-training, the model utilizes 100,000 hours of UMI trajectory data spanning over 1,700 scenarios, including homes, commercial facilities, industrial sites, and outdoor environments. Since manual labeling was impractical, Xiaomi built an automated annotation pipeline using a powerful VLM; this pipeline divides trajectories into fixed-length segments and assigns linguistic descriptions of the scene state transitions for each segment.
Post-training employs a cross-embodiment dataset including over 7,200 hours of proprietary robot data collected in actual residences (such as tidying sofas, sorting shoe racks, and clearing tableware), filtered open-source data, and high-quality UMI data.
During pre-training, scaling behavior was observed where action errors during validation decreased steadily as data volume and model size increased. Evaluations after post-training showed that in unseen environments and with unseen objects, the success rate of real robots increased predictably as pre-training data volume and model size grew, demonstrating that the scale of pre-training transfers to real-world robot performance.
Source: Xiaomi-Robotics-1 (HN 528pt, 325 comments) (HN Search (backfill), 2026-07-20)