A research team from Carnegie Mellon, MIT, NYU, and Stanford has developed Ataraxos, an AI that has achieved superhuman performance in Stratego, a classic board game characterized by massive amounts of hidden information. In a 20-game series, Ataraxos defeated Pim Niemeijer, widely considered the best Stratego player of all time, with 15 wins, 1 loss, and 4 draws.
Ataraxos AI Achieves Superhuman Performance in Stratego with Low Compute Cost
Unlike previous efforts like DeepMind's DeepNash, which required millions of dollars in compute, Ataraxos was trained using approximately 16 GPUs for one week and four additional GPUs for four days. The total training cost is estimated to be less than $8,000 at 2025 prices. This efficiency was achieved through a custom CUDA-accelerated simulator capable of executing millions of moves per second.
Ataraxos utilizes a design combining self-play reinforcement learning and test-time search. A key innovation is its "belief model"—a second neural network trained to estimate the opponent's hidden pieces based on their movements. This allows the AI to perform look-ahead searches by sampling plausible board configurations, overcoming the massive search space that previously limited AI performance in imperfect-information games.
The researchers demonstrated the scalability of this approach by applying similar techniques to other games. They successfully developed state-of-the-art AIs for Barrage Stratego, Hanabi (a cooperative card game), and dou dizhu (a popular Chinese card game), achieving high performance across adversarial, cooperative, and team-based settings.
Sources
- With most information hidden, the game Stratego (Ars Technica AI, 2026-10-01)
- Nature