SpaceXAI rolls out Grok 4.7 with higher scores at same price

SpaceXAI has just rolled out Grok 4.7, promising stronger reasoning, better self-checking, and its “strongest safeguards to date” without raising price or slowing down. Benchmarks show gains over Grok 4.6, but the comparison setup is drawing criticism.

SpaceXAI rolls out Grok 4.7 with higher scores at same price

TL;DR

  • Grok 4.7 announced: Positioned as better than Grok 4.6 at the same price and speed
  • Availability: In Cursor, Grok Build, and the Grok API
  • Claimed changes: Longer work on hard tasks, more careful checking, strongest safeguards to date
  • Benchmarks: Grok 4.7 xHigh exceeds Grok 4.6 High across listed tests; not top vs competitors
  • Pricing: $2/M input tokens, $6/M output tokens (both Grok 4.7 xHigh and Grok 4.6 High)
  • Criticism and limits: Comparisons use 4.7 xHigh vs 4.6 High; SuperGrok Heavy at 98% used, reset Sep 23, 2026

SpaceXAI has announced Grok 4.7, describing it as an improvement over Grok 4.6 at the same price and speed. The model is available in Cursor, Grok Build, and the Grok API.

The company claims that Grok 4.7 “works longer on difficult tasks,” checks its work more carefully, and includes its “strongest safeguards to date.” It also shared a comparison of the two models building an open-world city game.

Benchmark results

SpaceXAI’s comparison materials list Grok 4.7 xHigh alongside Grok 4.6 High and several other models:

BenchmarkGrok 4.7 xHighGrok 4.6 HighGPT-5.6 Sol MaxFable 5.1 Max
CursorBench 4.046.3%40.4%41.7%51.8%
DeepSWE v1.171.0%*65.2%72.7%70.0%
EEBench64.0%53.0%39.4%56.4%
AA-Briefcase v1.11657154614871678
Terminal-Bench 4.038.0%20.3%37.3%57.9%
Harvey Legal Agent19.6%15.8%2.5%6.7%
HealthBench Pro56.7%48.5%60.5%62.1%

*High Effort.

The figures show Grok 4.7 ahead of Grok 4.6 across the listed benchmarks, although it does not lead every comparison. Fable 5.1 Max records the highest scores on CursorBench, AA-Briefcase, Terminal-Bench, and HealthBench Pro, while GPT-5.6 Sol Max leads DeepSWE v1.1.

A CursorBench chart also compares model scores with average cost per task. The listed prices for Grok 4.7 xHigh and Grok 4.6 High are identical at $2 per million input tokens and $6 per million output tokens. GPT-5.6 Sol Max is listed at $4 per million input tokens and $20 per million output tokens, while Fable 5.1 Max costs $10 per million input tokens and $50 per million output tokens.

Comparison draws criticism

The benchmark presentation has prompted questions about the evaluation setup. Several replies pointed out that the materials compare Grok 4.7 xHigh with Grok 4.6 High rather than Grok 4.6 xHigh, making the generational comparison less direct.

Another reply questioned the omission of Astra from the charts. The supplied material does not include a response from SpaceXAI addressing either criticism.

Grok 4.7’s rollout also coincides with usage limits for some of the company’s products. A usage page records the weekly SuperGrok Heavy allowance at 98% used, including Grok Build at 98%, with a reset scheduled for September 23, 2026, at 22:37. A separate Weekly Grok Bot Limit is listed at 0% used and scheduled to reset at 22:52 the same day. Additional usage credits are listed at US$0.00, alongside an option to buy more quota.

Source: SpaceXAI on X

Continue the conversation on Slack

Did this article spark your interest? Join our community of experts and enthusiasts to dive deeper, ask questions, and share your ideas.

Join our community