English

Product LaunchesCarnegie Mellon UniversityMassachusetts Institute of TechnologyNew York UniversityStanford UniversityAtaraxos

Ataraxos AI Achieves Superhuman Performance in Stratego with Low Compute Cost

A research team from Carnegie Mellon University, MIT, New York University, and Stanford University has developed Ataraxos, an AI capable of defeating the world's top Stratego players. In a 20-game series against Pim Niemeijer, arguably the greatest Stratego player of all time, the AI achieved 15 wins, 1 loss, and 4 draws.

Stratego is an "imperfect-information" game, meaning players cannot see the identities of their opponent's pieces until they engage in battle. This creates a massive state space and requires complex bluffing and long-term strategic planning. Unlike perfect-information games like Chess or Go, where all information is visible, the hidden nature of Stratego has historically made it difficult for AI to achieve superhuman performance.

Ataraxos overcomes these challenges through a novel design that combines self-play reinforcement learning with test-time search. The system utilizes two interdependent Transformer-based networks: a set-up network to determine initial piece arrangements and a move network to select actions. A key innovation is the inclusion of a "belief model," a second neural network trained to estimate the likely identities of an opponent's hidden pieces based on their movement patterns. This allows the AI to perform a search by sampling plausible game states and selecting the best move across them.

Notably, Ataraxos achieved these results with extreme efficiency. While previous AI efforts, such as DeepMind's DeepNash, required massive computational resources—estimated at $3 million to $4.5 million for training—Ataraxos was trained using just 16 NVIDIA H100 GPUs for one week and additional GPUs for the belief model, costing less than $8,000.

The researchers demonstrated the versatility of their approach by applying the same techniques to other games, achieving state-of-the-art results in Hanabi (a cooperative card game) and dou dizhu (a Chinese card game). The team suggests that the success of this design pattern could eventually be applied to real-world strategic decision-making, such as negotiations or military simulations, which also involve significant hidden information.

Sources

  1. With most information hidden, the game Stratego had stumped AI–until now (Hacker News Frontpage, 2026-10-02)
  2. Nature