xAI has made Grok 4.6 available via Cursor and Grok Build. It is also provided through APIs and partners such as OpenRouter, Vercel, and Cloudflare. The company stated that it will provide double the usage limits within Grok Build and Cursor during the first week.
Regarding performance, xAI claims to have achieved frontier intelligence in benchmarks for multi-agent coding and knowledge work. According to the Artificial Analysis Intelligence Index, the model recorded scores equivalent to GPT-5.6 Sol. Figures for competitors were cited from system cards and leaderboards published by each respective developer.
In the training process, xAI conducted longer supplementary training than it did for Grok 4.5. This involved curated model-generated data for reasoning and advanced technical concepts, high-quality engineering data, and improved optimizers and training recipes. During the SFT stage, Grok 4.5 was used to regenerate trajectories in fields such as STEM and software engineering, with problematic traces excluded via model-based checks.
Furthermore, in addition to knowledge work and general coding, Grok 4.6 was trained on a wide range of agent RL tasks, including domain-specific environments such as kernel optimization, web development, and CAD. It is reportedly particularly excellent at converting broad product ideas into working first drafts, consistently handling everything from researching unfamiliar fields to structuring applications and implementing core interactions. In long trajectories, self-testing behavior—where the model verifies its own work before proceeding to the next step—was also observed.
For visual and interactive projects, Grok 4.6 provides a stronger first pass than Grok 4.5. When given a concrete product idea, it can establish the application's structure and visual language in a single pass. xAI stated that this is especially useful for projects where the fastest path is to start by building something tangible and then iterating in a loop.
Safety measures have been improved and adjusted to match the model's capabilities. They are designed to maximize utility and security for legitimate use cases, such as vulnerability patching and accelerating engineering design cycles. The test suite for pre-deployment capability and safety calibration was the largest to date, and extensive post-deployment and third-party testing were also conducted.
Pricing starts at $2 per million input tokens and $6 per million output tokens. A high-speed variant with double the pricing is also available.
Source: Grok 4.6 (HN 632pt, 616 comments) (HN Search (backfill), 2026-08-13)