English

Model ReleasesPricing & LimitsSpaceX AIGrok 4.6

Grok 4.6 Scores 61 on Intelligence Index, Achieving Performance Near Claude Opus 5 at a Similar Price Point

This article is a translation. Read the Japanese original

SpaceXAI has released the performance evaluation results for Grok 4.6. It recorded 61 points on the Artificial Analysis Intelligence Index. This represents an increase of 5 points over Grok 4.5 and a 23-point increase compared to Grok 4.3.

This score is on par with GPT-5.6 Sol (max). It is positioned just below Claude Opus 5 (max, 63 points) and Claude Fable 5 (max with fallback, 62 points). While Grok 4.5 was released only about a month ago, this latest update has further pushed its performance.

Evaluation for agent-based tasks is also high. Its Elo on GDPval-AA v2 is 1753, the second-highest level after Claude Opus 5. There is no statistical difference from Claude Fable 5 or Qwen3.8 Max, as their confidence intervals overlap. The cost per task is $0.84, which is the same as Kimi K3, though its intelligence index is slightly higher.

Pricing remains unchanged from the previous model. The input/output prices per 1M tokens are maintained at $2/$6. This is more than 60% lower than Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30). Maintaining prices for a frontier model is reported to be unusual.

On AA-Briefcase, a private benchmark for long-term agent tasks, it recorded an Elo of 1577. This result places it in the same class as Fable 5. Efficiency is also a key characteristic; the average number of turns is estimated at approximately 53, with input token counts around 500 million. In contrast, Claude Opus 5 (max) required approximately 103 turns and about 2 billion tokens.


Sources: Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index(HN 343pt・417コメント) (HN Search (backfill), 2026-08-13)