Published: 22 September 2026
SpaceXAI released Grok 4.7 on 21 September 2026, a model aimed at coding and knowledge work, priced at $2 per million input tokens and $6 per million output tokens. That is unchanged from Grok 4.6. It scores 46.3 percent on CursorBench 4.0 against 40.4 percent for Grok 4.6, and it carries a 500,000 token context window.
What is new in Grok 4.7?
SpaceXAI says Grok 4.7 sits on a new and larger base model, with a longer reinforcement learning run on harder problems, including tasks that take many hours to finish. The company credits the gains to better self-verification and better handling of long context. It takes text and images in and writes text out, with a knowledge cutoff of May 2026.
What you can use today:
- Access: the xAI API, Grok Build, Cursor, OpenRouter, Vercel and Cloudflare, live on announcement
- Free trial access through Grok Build
- Native Grok Bot harness support for agent workflows
- A Fast variant at $4 and $12 per million tokens for double the output speed
- Context window: 500,000 tokens
Grok 4.7 benchmarks against Claude Fable 5.1 and GPT-5.6 Sol
SpaceXAI published a seven-row comparison. Below is Grok 4.7 next to the strongest rival in each row of that same table.
| What is measured | Grok 4.7 | Best rival in SpaceXAI’s own table |
|---|---|---|
| CursorBench 4.0 (coding inside an editor) | 46.3% | 51.8% (Claude Fable 5.1) |
| DeepSWE v1.1 (agentic bug fixing) | 71.0% | 72.7% (GPT-5.6 Sol) |
| Terminal-Bench 4.0 (command line work) | 38.0% | 57.9% (Claude Fable 5.1) |
| EEBench (not defined in the announcement) | 64.0% | 56.4% (Claude Fable 5.1) |
| Harvey Legal Agent (legal agent tasks) | 19.6% | 6.7% (Claude Fable 5.1) |
| HealthBench Professional (clinical questions) | 56.7% | 62.1% (Claude Fable 5.1) |
These are SpaceXAI’s own evaluations and the post does not say where the rival figures came from or which harness produced them. Claude Fable 5.1 tops four of the seven rows. Grok 4.7 wins two, and on safety evaluations SpaceXAI reports 62.4 percent on the LatchBio biosafety benchmark and a 3.3 percent pass-through rate on risky dual-use prompts in HackerBench v0.3.
Why Grok 4.7 can cost more than its price suggests
The price per token did not move. The number of tokens did. Artificial Analysis measured Grok 4.7 in xHigh mode using about 81,000 output tokens per Intelligence Index task, against roughly 36,000 for Grok 4.6 in High mode. That is 125 percent more output for the same job.
Work that through and the headline falls apart. On the same index, a task runs to about $3.74 on Grok 4.7 against about $1.99 on GPT-5.6 Sol, even though Grok is cheaper per token. SpaceXAI does not publish token consumption anywhere in the announcement, and it is the figure that decides your bill. The Fast variant makes this sharper, not softer: double the speed at double the price on a model that already emits far more tokens than its predecessor.
What this means
Worth testing if you already run Grok, and worth skipping if you are choosing fresh on cost. A platform team migrating from Grok 4.6 gets a real jump on the command line, from 20.3 to 38.0 percent on Terminal-Bench 4.0, at no change in list price, and that is a clean upgrade. The exception is legal work, where 19.6 percent on the Harvey Legal Agent benchmark against 6.7 percent for Claude Fable 5.1 is the widest lead SpaceXAI has anywhere, and the one result that would make us test this model on purpose rather than by default.
For anyone else the comparison to make is cost per finished task, not cost per million tokens. Grok 4.7 loses four of seven benchmarks to Claude Fable 5.1 in SpaceXAI’s own table and costs nearly twice as much as GPT-5.6 Sol to complete the same index task. A $2 and $6 price card reads well in a procurement meeting and does not survive contact with a reasoning model that thinks twice as long. Run your ten hardest real tickets through it, measure the output tokens, and decide on that number.
For more information, visit the official announcement of Grok 4.7 on the SpaceXAI blog.