Software engineer Ben Swerdlow has released "Brood War Bench," a new benchmark designed to evaluate the ability of general-purpose AI agents to play the real-time strategy game StarCraft: Brood War. Unlike specialized AI like DeepMind's AlphaStar, which was trained specifically for StarCraft II, this benchmark tests models typically used for text and code generation, measuring how well they can autonomously handle real-time gameplay.
Brood War Bench Benchmark Evaluates General-Purpose AI Agents in StarCraft: Brood War
The benchmark involved a round-robin tournament with 19 different configurations of models and reasoning effort levels, including agents from OpenAI, Anthropic, and xAI. Among the results, Codex Astra achieved a perfect record of 18 wins and 0 losses. However, Swerdlow noted that even the top-performing models did not exceed the skill level of a beginner, suggesting that strategic complexity and execution remain significant hurdles.
The testing revealed distinct behavioral patterns among the models. Codex-based agents often focused on disruptive tactics, such as sending Probe units to harass enemy workers, but struggled with sustained production and coordination between sub-agents. Claude Fable showed a tendency toward long-term economic growth and technological advancement, successfully producing advanced units like Mutalisks in some matches.
In contrast, some models struggled with the fundamental requirements of real-time interaction. Grok 4.6, for instance, frequently failed to maintain a consistent control loop, sometimes issuing very few commands despite significant reasoning time. Swerdlow observed that the difficulty lies not just in strategy, but in the ability to observe changing game states and act promptly. He suggested that a human player using a basic "photon rush" rush tactic could likely defeat all participating agents.
Sources
- AIに「StarCraft」をリアルタイムで戦わせる「Brood War Bench」が登場、全モデルが初心者レベルながらCodex Astraは18戦全勝 (GIGAZINE, 2026-09-25)
- Brood War Bench