Microsoft Research has announced the addition of offloaded physical AI inference capabilities to its Physical AI Toolchain. This new feature allows robotics workloads to be distributed between a robot's local compute, edge GPUs, or the cloud using Kubernetes to manage containerized workloads.
Product LaunchesMicrosoft ResearchRho
Microsoft Research Introduces Offloaded Inference Capabilities for Physical AI Robotics
A recent systematic study conducted by the researchers focused on mobile robotic manipulation to evaluate the impact of inference location. The study found that performing inference on-device via onboard GPUs presents significant constraints as AI models grow in size. The researchers observed that smaller GPUs could not adequately handle mobile manipulation stacks, with some mapping and planning tasks slowing by up to 383% compared to an A100 GPU. Additionally, the study noted that lightweight GPUs led to a 30% decrease in timely obstacle detection for navigation and a 50% drop in accuracy for Vision-Language-Action (VLA) models.
Beyond computational performance, the study highlighted the impact on power consumption. Comparing an onboard GPU to a Raspberry Pi-5 (with data shipped to an offloaded GPU), the researchers found that large onboard GPUs, such as the Jetson Thor, could increase battery drain by up to 160%, potentially reducing operation time by several hours.
The new offloading capability is integrated into the Physical AI Toolchain, an open-source framework that combines Microsoft Azure services with NVIDIA’s physical AI stack. To demonstrate the capability, Microsoft provided example projects for offloading inference on the SO-101 and UR10e robots. The company also demonstrated the offloading of its Rho model—designed for dual-arm robots—to a Jetson Thor GPU controlling a Mobile Aloha robot.
Sources
- Offloaded inference for real-world physical AI robotics (Microsoft Research, 2026-09-23)
- Microsoft Research Forum