Updated 19 hours ago
Grok 4.7 improves agent benchmarks, but task costs complicate the pitch

AI model benchmarks

Grok 4.7 improves agent benchmarks, but task costs complicate the pitch

Grok 4.7 posts stronger long‑horizon agent results at unchanged token rates, while independent testing finds higher usage, elapsed time and pay‑per‑token API cost per coding task.

SpaceXAI [released Grok 4.7](https://x.ai/news/grok‑4‑7) on September 21 with an appealing headline: more capable coding and knowledge work at the same token prices as Grok 4.6. The model is already available through the public xAI API as `grok‑4.7`, as well as in Cursor and Grok Build, so the upgrade decision is immediate rather than theoretical. The harder question is what “same price” means once the model starts working. The first independent results point to a real improvement on long‑running agent tasks. They also show that Grok 4.7 can spend much more time and many more tokens reaching those results. For teams paying by usage or waiting on autonomous jobs, the model's list price, benchmark score and benchmark pay‑per‑token API cost now tell three different parts of the story.

The release is aimed at work that takes longer to finish

[SpaceXAI says](https://x.ai/news/grok‑4‑7) Grok 4.7 uses a larger base model than Grok 4.6 and received a longer reinforcement‑learning run weighted toward tasks that take hours to complete. The company also claims better self‑verification and long‑context handling, positioning the model for coding agents, document production and other work where a system must plan, use tools and recover from mistakes over many steps. Those are first‑party descriptions, but they identify the workload the release is supposed to improve rather than promising an across‑the‑board jump. The [xAI developer documentation](https://docs.x.ai/developers/grok‑4‑7) lists a 500,000‑token context window, text and image input, text output and four reasoning settings from low through xhigh. The public model ID is `grok‑4.7`. SpaceXAI's launch page says the standard model keeps Grok 4.6's starting rates of $2 per million input tokens and $6 per million output tokens, while a faster serving option costs twice as much. That faster option has a meaningful availability limit. SpaceXAI says Grok 4.7 Fast is the same model on faster infrastructure, but it is available only in Cursor and Grok Build, not through the public xAI API. API users can access a US regional endpoint, where the documentation says token usage carries a 10% premium, but they cannot select the Fast tier there. Cursor presents Grok 4.7 as a model for difficult, long‑running work. Its [Grok 4.7 documentation](https://prod.cursor.com/docs/models/grok‑4‑7) lists a standard 256,000‑token context window and a 500,000‑token maximum, with high as the default reasoning effort. That distinction matters because more reasoning effort may be valuable for a stubborn migration or repository‑scale investigation and excessive for a routine edit.

Independent testing finds its largest gains in agentic work

[Artificial Analysis evaluated Grok 4.7](https://artificialanalysis.ai/articles/benchmarking‑grok‑4‑7) at xhigh effort and scored it 46 on its Intelligence Index, two points above Grok 4.6 at high effort. The more pronounced change appeared in the firm's Coding Agent Index, where Grok Build with Grok 4.7 scored 56, up nine points from Grok 4.6 at xhigh. In the snapshot published on launch day, that placed the combination fourth among models running in their native coding‑agent harnesses, behind Claude Fable 5.1, GPT‑6 Astra and Claude Opus 5. The coding index is not one synthetic question set. Artificial Analysis says its [current methodology](https://artificialanalysis.ai/methodology/coding‑agents‑benchmarking) equally weights DeepSWE v1.1, Terminal‑Bench 4.0 and SWE‑Atlas‑QnA, covering implementation, terminal work and technical repository questions. Grok 4.7 with Grok Build improved across all three components in the published comparison: 73% versus 65% on DeepSWE, 33% versus 18% on Terminal‑Bench, and 63% versus 58% on SWE‑Atlas‑QnA. The gain was narrower outside those agentic workloads. Artificial Analysis reported that Grok 4.7 broadly matched Grok 4.6 on the other Intelligence Index tasks, with improvements on Terminal‑Bench and GDP.pdf and regressions on AA‑LCR and AutomationBench‑AA. Its separate factuality measure compared Grok 4.7 at xhigh with Grok 4.6 at high: the hallucination rate was 29% versus 34%, while accuracy was nearly unchanged at 47% versus 48%. That combination supports a focused reading of the release: the clearest evidence is better execution on long‑horizon tasks, not a universal capability leap. There is also an important harness caveat. The Intelligence Index standardizes how models are tested, while the Coding Agent Index uses each model inside a native agent; Grok 4.7 was paired with Grok Build. The coding result therefore measures the model‑and‑agent system that a user might actually run, but it cannot isolate how much of the improvement came from the model, the harness or their interaction.

Unchanged token rates produced higher measured task costs

SpaceXAI's $2 input and $6 output rates are unchanged from Grok 4.6, yet equal per‑token prices do not guarantee equal bills. Artificial Analysis reports that Grok 4.7 at xhigh used about 81,000 output tokens per Intelligence Index task, more than twice the roughly 38,000 used by Grok 4.6 at xhigh. At the $6 output rate, those counts alone correspond to about $0.49 and $0.23 in output charges respectively, before input tokens, cache behavior, tools, retries or any platform fees are counted. The [independent coding‑agent comparison](https://artificialanalysis.ai/agents/coding‑agents/comparisons/grok‑build‑vs‑muse‑code) shows the same tension at the task level. Artificial Analysis currently lists Grok Build with Grok 4.7 xhigh at an average $8.82 in pay‑per‑token API cost and 39.2 minutes per task, compared with $3.57 and 19.5 minutes for Grok 4.6 xhigh. Its cost metric applies provider token prices to the benchmark's measured usage; it is not a consumer subscription or an all‑in production estimate, and it excludes infrastructure, engineering and human supervision. Within that scope, Grok 4.7 measured about 2.47 times the cost and twice the wall time for a nine‑point index gain in this benchmark environment. Those figures do not establish what Grok 4.7 will cost on a particular repository. They do establish why the vendor's unchanged rate card cannot settle the upgrade question. A model that succeeds on a task its predecessor fails may be cheaper than repeated attempts, manual repair or escalation to a more expensive model; a model that produces the same acceptable patch with twice the work may be a poor trade even at an attractive token rate. Artificial Analysis measured output generation at approximately 188 tokens per second on long prompts, while an average Intelligence Index task took about 7.1 minutes. Both can be true because token throughput measures one part of an agentic run, while total time also includes reasoning and the steps surrounding generation. Teams evaluating “fast” claims should measure elapsed job time rather than infer it from output speed alone.

Cursor and the public API cross long‑context pricing at different points

The release also has two different billing thresholds that can easily be collapsed into one. SpaceXAI's [September 21 API release note](https://docs.x.ai/developers/release‑notes) says public API requests cost $2 per million input tokens, $0.50 per million cached input tokens and $6 per million output tokens below a 200,000‑token prompt. Above that threshold, the rates double to $4, $1 and $12 respectively. Cursor documents a 256,000‑token standard window and doubles standard rates when input exceeds 256,000 tokens, up to 500,000. It lists standard on‑demand usage at $2/$0.50/$6, Fast at $4/$1/$12, standard long‑context usage at $4/$1/$12, and Fast long‑context usage at $6/$1.50/$18 per million input, cached‑input and output tokens. The same model name can therefore produce a different bill depending on where it runs, how much context is sent and whether the faster tier is selected. That makes context management part of model selection. A large repository dump that crosses a billing threshold can increase the rate applied to a request, while disciplined retrieval, caching and compaction may keep the useful context smaller. SpaceXAI itself recommends a prompt cache key for reliable cache affinity and context compaction for long agent loops, practical details that matter more to recurring cost than the headline context maximum.

What teams should test before moving difficult work to Grok 4.7

A useful evaluation should begin with a fixed set of the team's own tasks: a multi‑file bug, a migration, a repository question, a document or spreadsheet deliverable, and one ordinary edit that should not require an expensive agent loop. Run the same repository snapshot and acceptance checks with Grok 4.6 and 4.7, keep the harness and effort setting constant, and record completion, human corrections, retries, total tokens, cache hits, elapsed time and the final charge. Repeating each task is more informative than treating one successful demo as a stable rate. Effort settings deserve their own comparison because the published independent results use xhigh, while Cursor defaults to high. Long‑context tests should also be separated from normal‑context work so the higher pricing tier does not distort an otherwise routine comparison. The evidence available on launch day makes Grok 4.7 a credible candidate for tasks where persistence and tool use determine success, but it also makes the cost of that persistence a result teams need to measure rather than assume.

Sources

Share this article

PostShare

Related News