xAI has announced Grok 4 Fast, an inference model with superior cost efficiency. The company explained that it maximized intelligence density through large-scale reinforcement learning.

Grok 4 Fast demonstrates performance exceeding Grok 3 Mini in inference benchmarks. Simultaneously, it has reduced average thinking token usage by 40%.

Due to this efficiency, the price to achieve performance equivalent to Grok 4 is reduced by 98%. A review by the independent evaluation firm Artificial Analysis also states that the model exhibits one of the highest cost-to-intelligence ratios among publicly available models.

Furthermore, Grok 4 Fast was trained using reinforcement learning for tool use. It excels in the ability to appropriately call tools such as code execution and web browsing as needed.

In the LMArena Search Arena, grok-4-fast-search—provided under the codename "menlo"—earned 1163 Elo, ranking 1st. This result surpasses o3-search by 17 points.

In the Text Arena, grok-4-fast, under the codename "tahoe," is positioned 8th. While other models in the same size range remain below 18th place, it exhibits performance equivalent to the comparative Grok 4.

Additionally, the model employs an architecture that integrates reasoning and non-reasoning modes within a single model. The company states that this reduces end-to-end latency and token costs.


Source: Grok 4 Fast (HN 96pt, 76 comments) (HN Search (backfill), 2025-09-20)