What shipped
SpaceXAI says Grok 4.7 is built on a larger base model than Grok 4.6 and went through a longer reinforcement-learning phase focused on tasks requiring several hours of work. The company also improved long-context handling and trained the model to work better with the recently released AI-agent-oriented Grok Bot interface.
That positioning targets teams using models as agents rather than chat interfaces. A model that can continue a long coding session, assemble a document, work in a terminal or complete a multistep business process has different operating requirements from one designed mainly for individual prompts.
Grok 4.7 is available through:
Cursor, which SpaceXAI recently acquired, offers Grok 4.7 Fast at twice the standard price: $4/$12 per 1 million input and output tokens. It is included by default in the Pro and higher tiers.
The fast version supports input requests larger than 256,000 tokens and a context window of up to 500,000 tokens. Cursor does not publish a tokens-per-second figure or promise exactly twice the throughput. For inputs above 256,000 tokens, standard Grok 4.7 costs $4/$12, while the fast version costs $6/$18.
SpaceXAI also offers Grok 4.7 Fast through its own platform at $4/$12 per 1 million input and output tokens. The company has not disclosed its tokens-per-second figure, and independent throughput measurements were unavailable at publication.
Better scores, familiar limits
On several widely used third-party benchmarks, Grok 4.7 clearly improves on its predecessor. It still trails the latest models from OpenAI, Anthropic and Google on some tests.
Developer and author Dan McAteer called Grok 4.7’s Terminal-Bench result “terrible.” That benchmark measures terminal work and showed the model’s largest gain over Grok 4.6.
Artificial Analysis measured Grok 4.7 xHigh at about 26% on Terminal-Bench. OpenAI GPT-6 Astra xHigh scored 59.6%, while Anthropic Claude Opus 5 at its maximum effort level scored about 49%.
Other results include:
Some observers saw the release more favorably. The popular AI-news account @haider1 called the CursorBench result “actually absurd.” In the comparison it shared, Grok 4.7 xHigh scored 46.3% at about $6.01 per task. Fable 5.1 Medium scored 46.8% at $7.05, while GPT-5.6 Sol Max scored 41.7% at $8.23. The account also pointed to Grok 4.7’s improvement over Grok 4.6 at roughly comparable cost.
The more important comparison, however, is not price per token. It is price per completed task.
Cheap tokens, expensive work
Artificial Analysis found that Grok 4.7 xHigh used about 81,000 output tokens for one Intelligence Index task. Grok 4.6 High used 36,000, and GPT-6 Astra Max used 27,000.
That means Grok 4.7 used 125% more output tokens than Grok 4.6 and 196% more than GPT-6 Astra.
Artificial Analysis estimated the resulting cost as follows:
The methodology includes input, cached, reasoning and output tokens rather than comparing API list prices alone. That distinction matters: a model with a lower token price can cost more to complete the same job if it needs substantially more reasoning.
At $2/$6, a workload of 10 billion input tokens and 2 billion output tokens per month would cost about $32,000 per month, or $384,000 per year. At 50 billion input tokens and 10 billion output tokens, the bill would reach about $160,000 per month, or $1.92 million per year, excluding platform fees.
These are illustrative calculations, not estimates of a typical Grok deployment. They show why token efficiency becomes an infrastructure-budget issue at production scale.
My guess is that this is the real test of Grok 4.7’s pricing strategy. The model has enough benchmark improvement to justify a serious trial, but the trial should measure more than API rates:
Safety claims still need outside checks
SpaceXAI says Grok 4.7 has an entirely new set of protective mechanisms. The model scored 62.4% on the LatchBio biosafety benchmark. In HackerBench v0.3, a test for dangerous dual-use tasks, the company says it allowed only 3.3% of risky prompts through.
Those are SpaceXAI’s results, so they should be treated as company-reported figures until broader independent evaluations are available.
For engineering teams considering Grok 4.7 this quarter, the key question is not whether $2/$6 looks cheaper than a competitor’s price sheet. It is whether the model can complete real programming and knowledge-work tasks with fewer dollars, retries and human interventions. Grok 4.7’s benchmark gains make that experiment worth running; its appetite for tokens may determine whether the result is a cheaper system or merely a cheaper invoice line.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X