SpaceXAI released Grok 4.6 on 12 August 2026, and the headline number is that it scores 61 on the Artificial Analysis Intelligence Index. That puts it level with GPT-5.6 Sol and one point behind Claude Fable 5, in a field where Claude Opus 5 leads at 63 on maximum reasoning effort. The interesting part is the price. Grok 4.6 charges $2 per million input tokens and $6 per million output tokens, which Artificial Analysis measures as more than 60 percent below both of the models it now matches.

This is not a new foundation model. SpaceXAI kept the base from Grok 4.5 and spent the improvement on post-training, which reframes the whole upgrade question. Below are the benchmark figures with their version numbers intact, plus the full token rates including the 200,000-token threshold that doubles your bill. After that, an honest read on whether Grok 4.5 users should move, and how to run the model on a Mac without a developer setup.

The Key Takeaways

  • Released: 12 August 2026 by SpaceXAI, the company formerly known as xAI.
  • Intelligence: 61 on the Artificial Analysis Intelligence Index, tied with GPT-5.6 Sol, behind Claude Opus 5 at 63 and Claude Fable 5 at 62.
  • Price: $2 in, $6 out per million tokens, unchanged from Grok 4.5 and roughly a third of what Claude Opus 5 charges at $5/$25.
  • The trap: cross 200,000 prompt tokens and SpaceXAI re-bills the entire request at $4/$1/$12, not just the overage.
  • Efficiency: Artificial Analysis clocks it at $0.84 per task, finishing long jobs in about 53 turns against roughly 103 for Claude Opus 5.

What Grok 4.6 Actually Changes

Od vydavatele

Každý AI model v jedné aplikaci

Fello AI přináší GPT-5.6, Claude 5, Gemini 3.6, Grok 4.5 a další v jedné nativní aplikaci pro Mac a iPhone.

Stáhnout hned!

The framing SpaceXAI used at launch is worth taking literally. Grok 4.6 is a refinement of Grok 4.5 aimed at long-running agents, coding, and what the company calls more ambitious interactive and visual work. The company held the foundation constant and bought the gains with a longer supplemental training run, regenerated fine-tuning trajectories, and reinforcement learning inside agentic environments.

SpaceXAI has published no new architecture for this release. No parameter count, no new base model name. Anyone quoting one is reading it across from the previous generation rather than from the announcement.

What did change is behaviour over long sessions. Cursor, which shipped the model on day one, reported that Grok 4.6 does more self-testing and verification, checking its own work before moving to the next step. The same report notes stronger first passes on visual and interactive projects than Grok 4.5 produced. That is the kind of improvement that shows up in agent work and barely registers in a single-turn chat.

The specs themselves are unchanged in the places that matter for compatibility.

SpecDetail
API model namegrok-4.6
Released12 August 2026
Context window500,000 tokens
Input modalitiestext and images
Reasoning effortlow, medium, high (default), xhigh
Base modelsame foundation as Grok 4.5, no new architecture published
Fast variantavailable at twice the standard price

Availability was broad from the first day. The model went live in the SpaceXAI API, in Cursor and Grok Build, and through partners including OpenRouter, Vercel and Cloudflare. Cursor also offered double the included usage inside Cursor and Grok Build for the launch week.

Grok 4.6 Benchmarks: What the Numbers Say

Two sets of numbers exist for this model, and they measure different things. Keeping them apart is the only way to read the release honestly.

SpaceXAI's Own Launch Table

These are the figures from the official Grok 4.6 announcement, all measured at high reasoning effort. SpaceXAI reports Grok 4.6 improving on Grok 4.5 across every evaluation it lists, which is a vendor claim about a vendor benchmark suite and should be read as one.

BenchmarkGrok 4.6 (high)What it measures
AA Intelligence Index61composite of nine benchmarks
GDPval-AA v21753 Elosustained agentic knowledge work
CursorBench v3.269.9%real editor sessions
DeepSWE v1.165.9%software engineering tasks
FrontierCode v1.161.3%hard coding problems
APEX-Agents57.5%multi-step agent execution
APEX-SWE56.4%agentic software work
Terminal-Bench v3.026%command-line task completion
AA-Briefcase1577 Elobusiness and analyst tasks
Harvey LAB15.8%legal agent work

That Terminal-Bench figure needs a warning label, because a lot of coverage has already mangled it. Terminal-Bench v3.0 is a different and much harder test than v2.1. Grok 4.6 scoring 26% on v3.0 and 88.4% on v2.1 are both true statements about the same model. Any article printing one number without its version is not telling you anything useful.

Independent Scores from Artificial Analysis

The independent analysis from Artificial Analysis is the number to lean on, because the lab runs every model through the same harness. Grok 4.6 scores 61 on its Intelligence Index. Claude Opus 5 leads the board at 63, Claude Fable 5 sits at 62, GPT-5.6 Sol ties Grok 4.6 at 61, and Kimi K3 lands just behind. Joint third place on the frontier, from a model released as a post-training refresh, is a genuine result.

The agentic figures are stronger than the composite. On GDPval-AA v2, Grok 4.6 posts an Elo of 1753, behind only Claude Opus 5. It reaches 50.7% on the tau3-Banking financial reasoning task and 88.4% on Terminal-Bench v2.1, which the lab describes as in line with the leading models.

Efficiency is where the release separates itself. Artificial Analysis measures Grok 4.6 at $0.84 per task, matching Kimi K3 while scoring higher on intelligence. On long-horizon work it finishes in roughly 53 turns and 0.5 billion input tokens, against about 103 turns and 2.0 billion input tokens for Claude Opus 5. Fewer turns and fewer tokens for comparable output is the practical argument for this model, more than any single benchmark row.

Grok 4.6 Pricing and the 200K Cliff

Start with the part nobody puts in a headline. Grok 4.6 has two price tiers, and which one you pay is decided by the size of your prompt.

Prompt sizeInput / 1MCached input / 1MOutput / 1M
Under 200,000 tokens$2.00$0.50$6.00
200,000 tokens and above$4.00$1.00$12.00

According to the SpaceXAI model documentation, once a prompt reaches 200,000 tokens, every token in that request bills at the higher rate. Not the overage. The whole thing.

Work the arithmetic on a 201,000-token prompt and the shape of it lands. At the standard rate that input costs about $0.40. Crossing the threshold by a single thousand tokens takes it to roughly $0.80, and the output on that same request doubles from $6 to $12 per million. A 500,000-token context window with a toll booth at 200,000 is a different product from one without.

The fix is unglamorous and effective: keep long-context jobs under the threshold by trimming what you paste in, or accept the doubled rate deliberately rather than discovering it on an invoice. For the full picture across plans and tiers rather than raw token rates, the complete breakdown of Grok pricing covers the subscription side.

Set against the field, the headline rates are the story. Artificial Analysis puts Grok 4.6 more than 60 percent below Claude Opus 5 at $5 in and $25 out, and below GPT-5.6 Sol at $5 and $30. Those are the two models Grok 4.6 is now measured against, and it undercuts both by a wide margin while scoring within two points of them.

Grok 4.6 vs Grok 4.5: Should You Move?

SpecGrok 4.5Grok 4.6
Released8 July 202612 August 2026
AA Intelligence Index5661
Terminal-Bench v3.015.7%26%
Context window500,000 tokens500,000 tokens
Input / output per 1M$2 / $6$2 / $6
Base modelV9 foundationsame foundation, post-training upgrade

Same context window, same headline price, five points of measured intelligence, and command-line task completion up from 15.7% to 26% on Terminal-Bench v3.0. There is no cost argument for staying on 4.5 and no compatibility argument either, since the model name is the only thing that changes in an API call.

The upgrade earns its keep most clearly in agent work, long coding sessions, and multi-step research where the model has to hold a thread across dozens of turns. That is precisely what the training run targeted, and the turn-count efficiency figures back it up. For single-turn questions and short chats, the difference is close to invisible, which is an honest thing to say about a release that moved five points on a composite index.

How to Use Grok 4.6 on a Mac

Every launch-day guide to this model assumes you have an API key and a terminal open. Most people do not, and the model is far more useful than that audience implies.

The routes that need no developer setup are the Grok app and grok.com, both covered by SpaceXAI's own subscription tiers, and third-party clients that connect over the API. If you want Grok alongside the other frontier models rather than in its own silo, a native Mac client is the cleaner answer. The Grok desktop client for macOS runs it as a proper Mac app, with no browser tab and no X subscription required. The same app carries GPT-5.6, Claude Opus 5 and Gemini, so switching models is a dropdown rather than a second account.

That last point matters more this month than usual. With four models inside two points of each other on the same index, and each one leading a different category, the practical setup is access to several of them. Our running comparison of the best AI models available right now tracks where each one currently leads.

The Verdict

Grok 4.6 is the best value on the frontier right now, and that is a narrower claim than being the best model. Claude Opus 5 still leads the Intelligence Index at 63 and still wins GDPval-AA v2. What Grok 4.6 does is arrive within two points of the leader, beat almost everything else on cost per task, and finish long agentic jobs in half the turns.

Move to it from Grok 4.5 today. The price is identical, the gains are measured by an independent lab rather than only by the vendor, and nothing in your setup breaks. Watch the 200,000-token threshold if you work with large documents, because that is the one place where the sticker price and the real price part company.

FAQ

Is Grok 4.6 free?

Not on the API, where it bills per token from the first request. SpaceXAI's consumer subscriptions include access to its current models, and Cursor offered double the included usage in Cursor and Grok Build during launch week. For what each plan includes, see the full Grok pricing breakdown.

Is Grok 4.6 better than GPT-5.6?

They tie. Both score 61 on the Artificial Analysis Intelligence Index. Grok 4.6 costs $2 and $6 per million tokens against $5 and $30 for GPT-5.6 Sol, so at equal measured intelligence Grok 4.6 is far cheaper to run.

Is Grok 4.6 a new base model?

No. SpaceXAI kept the foundation from Grok 4.5 and improved the model through additional post-training, including reinforcement learning in agentic environments. No new architecture or parameter count was published with the release.

How big is the Grok 4.6 context window?

500,000 tokens, the same as Grok 4.5. Prompts of 200,000 tokens or more bill at double the standard rate across the entire request, so the usable cheap window is smaller than the maximum window.

Where can I use Grok 4.6?

The SpaceXAI API, Cursor, Grok Build, and partners including OpenRouter, Vercel and Cloudflare. On a Mac, a native multi-model client runs it as a desktop app without a browser tab or an X subscription.