OpenAI has announced that its Astra model demonstrated extremely high performance on "ARC-AGI-3," a benchmark used to measure the intelligence of agents.
OpenAI's Astra Records High Scores on ARC-AGI-3 Benchmark
This article is a translation. Read the Japanese original
The ARC-AGI series aims to measure the gap between current AI and Artificial General Intelligence (AGI). ARC-AGI-3 specifically tests an agent's ability to inference goals and plan actions within unknown environments.
According to OpenAI, Astra (high) recorded a score of 99.9% in an environment using the "Provider Adapter harness," which the company stated is a state-of-the-art score. Additionally, Astra (max) recorded a score of 62.7% in an environment using the "Standard harness."
Astra manages information using its own notation, recording object coordinates, rules, and incomplete plans within the environment. Astra (max) solved tasks with fewer actions than humans at a 96.0% level.
In testing, Astra reduced the total cost across inference trials by selecting efficient actions.
Source: OpenAI's GPT-6 Astra on ARC-AGI-3 (Hacker News Frontpage, 2026-09-04)