Skip to main content

Grok 4.6 Arrives: Smarter, Cheaper, and Ready to Compete

The AI race just got a lot more interesting. xAI has officially rolled out Grok 4.6, and it's not just another incremental update. This model is making waves by delivering top-tier performance at a fraction of the cost of its rivals.

A Smart Move on the Intelligence Front

Grok 4.6 scored 61 on the AI Intelligence Index, tying with GPT-5.6 Sol (Max) and trailing only slightly behind Claude Opus 5 (63) and Claude Fable 5 (62). That puts it squarely in the global elite, rubbing shoulders with the best from OpenAI and Anthropic. But what's really turning heads is how it got there.

Building on Grok 4.5, the team at xAI focused on sharpening reasoning skills and advanced technical know-how. They used a clever trick: feeding the model its own high-quality outputs to retrain itself. By regenerating supervised fine-tuning data—covering everything from reasoning and agent tool use to STEM, software engineering, and knowledge work—they created a feedback loop that pushed performance to new heights.

Agent Benchmarks: A Huge Leap Forward

If there's one area where Grok 4.6 shines, it's in agent-related tasks. The numbers are impressive:

  • GDPval-AA v2: 1753 Elo, second only to Claude Opus 5
  • DeepSWE v1.1: jumped from 54% to 65.9%
  • APEX-Agents: climbed from 47.1% to 57.5%
  • Terminal-Bench v3.0: rose from 15.7% to 26%
  • CursorBench v3.2: hit 69.9%
  • FrontierCode v1.1: reached 61.3%

These aren't just marginal gains—they represent a significant leap in coding and agent capabilities. It's clear xAI is betting big on the "agent era," where AI handles complex, multi-step tasks with minimal human intervention.

The Cost-Efficiency Edge

Here's where things get really interesting. Grok 4.6 is priced at just $2 per million input tokens and $6 per million output tokens. Compare that to Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30), and you're looking at savings of over 60%. But the real kicker is how efficiently it works.

On the AA-Briefcase long-term knowledge work benchmark, Grok 4.6 completed tasks in an average of 53 rounds and about 500 million input tokens. Claude Opus 5? It needed roughly 103 rounds and a whopping 2 billion tokens. That means Grok 4.6 achieves the same results with less than a quarter of the resources. For businesses watching their bottom line, that's a game-changer.

Availability and Promotions

Grok 4.6 is already available through Cursor, Grok Build, API, and platforms like OpenRouter, Vercel, and Cloudflare. To sweeten the deal, xAI is offering double usage on Grok Build and Cursor during the first week. The context window remains a generous 500,000 tokens, so you can feed it plenty of data without hitting limits.

The Bottom Line

The message from xAI is clear: why pay more when you can get equal intelligence for less? With Grok 4.6, they've not only matched the big players on performance but also undercut them on price. It's a bold move that could reshape the competitive landscape. As the agent era unfolds, cost-effectiveness will be key, and Grok 4.6 is leading the charge.

Key Points

  • Grok 4.6 scores 61 on the AI Intelligence Index, tying with GPT-5.6 Sol (Max).
  • Significant improvements in agent benchmarks, with DeepSWE v1.1 rising from 54% to 65.9%.
  • Pricing is over 60% cheaper than Claude Opus 5 and GPT-5.6 Sol.
  • Achieves tasks with less than a quarter of the resources compared to Claude Opus 5.
  • Available now on multiple platforms, with promotional offers for the first week.