Two new open-source projects are introducing gaming-based benchmarks to evaluate AI model capabilities through different approaches.
New AI Benchmarking Projects Use Gaming to Measure Model Performance
Tiny AI Arena is a project that pits four randomly selected models against each other in turn-based battles on an 8×8 grid. The objective is to be the last survivor. Each model takes turns performing actions—such as moving, attacking adjacent enemies, or waiting—which consume Action Points (AP). The game includes environmental factors like obstacles and power-ups. A server acts as a referee to validate actions, and models can communicate through a 50-character global chat. Results are ranked on a leaderboard using an Elo rating system. According to the project's current rankings, Claude-Sonnet-5 holds the top position, followed by Claude-Fable-5.1 and Grok-4.6.
In a different approach, Pac-Man Bake-off evaluates how well models can recreate the game Pac-Man from a single prompt. The project compares different models and harnesses based on their ability to generate a playable HTML entry. Scoring for these entries is determined by factors such as controls, ghost behavior, Pac-Man getting stuck, the maze, and sound. According to a re-test of live games conducted on 2026-09-28 by Opus 5.5, scores were derived from 90-second automated play tests and maze audits.
Sources
- Show HN: Pac-Bench – How well can models one-shot a Pac-Man game? (Hacker News Frontpage, 2026-09-28)
- github.com/jonclegg/pacman-bakeoff
3 more sourcesHide sources
- AI同士をターン制のバトルゲームで対決させて能力を測る「Tiny AI Arena」 (GIGAZINE, 2026-09-29)
- GitHub
- 自分の脳が何B級のAIと同等性能なのか診断できるウェブサイト「HumanBench」が登場、選択問題で人間の力を試してどのAIに性格が似ているかも判定可能 (GIGAZINE, 2026-09-29)