The team developed a method that divides computational resources into an "islands" structure that coordinates asynchronously. This allows other parts of the system to continue training even if a failure occurs in one section.
Model ReleasesGoogle DeepMindDecoupled DiLoCo
DeepMind Releases Decoupled DiLoCo Distributed Training Architecture
This article is a translation. Read the Japanese original
Conventional synchronous approaches faced logistical difficulties in keeping thousands of chips in complete synchronization. Decoupled DiLoCo resolves this constraint.
This technology integrates "Pathways," which handles asynchronous data flow, with "DiLoCo," which achieves bandwidth reduction. It localizes the impact of chip failures and prevents total system shutdowns.
Verification using "Chaos engineering" demonstrated that training was maintained even after entire training units were lost. Once recovered, they are seamlessly reintegrated.
Tests with Gemma 4 models showed benchmark performance similar to conventional methods, while cluster availability during failures was improved.
Furthermore, the team successfully pre-trained a 12 billion parameter model spanning four regions in the United States. The wide-area network operated at 2–5 Gbps, a level achievable with existing internet connections.
This was more than 20 times faster than traditional synchronous methods. The design, which distributes communication over longer periods, contributed to this increase in speed.
Source: Decoupled DiLoCo: Resilient, Distributed AI Training at Scale (HN 49pt, 6 comments) (HN Search (backfill), 2026-04-28)