During his first week as Dev Rel Lead at OpenRouter, Jacky conducted an experiment pitting 11 different LLMs against each other in a custom-built 2D battle royale environment. Over 30 matches on a 400x400 meter map, each model recognized other participants as characters A through K and fought using weapons and items.

As a result of the experiment, Grok 4.1 Fast recorded 13 wins out of 30 matches, with a cost of $0.97 per win. In contrast, the runner-up, Claude Sonnet 4.6, secured 5 wins at a cost of $26.78 per win. The most affordable model outperformed the most expensive model in cost efficiency by a factor of 27.

GPT 5.4 recorded the highest number of kills at 38, but its total wins were limited to two. Furthermore, three models, including GPT 5.4-mini and DeepSeek 4 Flash, spent a combined total of $57 without achieving a single victory. Jacky stated that performance gaps that could not be predicted by existing benchmarks manifested as in-game behaviors.


Sources: A robot is sprinting towards you. Do you want it running on Claude or Grok? (HN 272pt, 21 comments) (HN Search (backfill), 2026-06-18)