SpaceXAI has announced Grok 4.6, describing it as a “frontier intelligence” model and a substantial improvement over Grok 4.5 at the same price. The company also claims that the new model is faster than comparable systems and can handle more difficult tasks.
Grok 4.6 is available through Grok Build, Cursor, Grok Bot and the API. Cursor and Grok Build will include twice the usual usage during the first week, according to SpaceXAI.
The API price is listed at $2 per million input tokens and $6 per million output tokens. SpaceXAI characterizes that as half the price of other frontier models, although the company’s own comparison material uses a separate cost figure: Grok 4.6 High is listed at $0.84, compared with $0.36 for Grok 4.5 High, $0.50 for GPT-5.5, $0.69 for Gemini 3.5 Flash and $0.81 for GPT-5.6 Sol.
Grok 4.6 leads Grok 4.5 across the listed tests
The benchmark comparison covers Grok 4.6 High, Grok 4.5 High, GPT-5.6 Sol Max and Fable 5 Max. It places Grok 4.6 High at 61 on the AA Intelligence Index, matching GPT-5.6 Sol Max and narrowly trailing Fable 5 Max at 62. Grok 4.5 High scores 56.
On GDPval-AA v2, Grok 4.6 High scores 1,753, ahead of GPT-5.6 Sol Max at 1,728 and Fable 5 Max at 1,741. Grok 4.5 High scores 1,526.
The results are less favorable on several software-engineering evaluations:
| Benchmark | Grok 4.6 High | Grok 4.5 High | GPT-5.6 Sol Max | Fable 5 Max |
|---|---|---|---|---|
| CursorBench v3.2 | 69.9% | 66.7% | 67.2% | 70.5% |
| DeepSWE v1.1 | 65.9% | 54% | 73% | 70% |
| FrontierCode v1.1 Extended | 61.3% | 56.6% | 60.6% | 64.9% |
| APEX-Agents | 57.5% | 47.1% | 56.7% | 59.2% |
| Terminal-Bench v3.0 | 26% | 15.7% | 34.6% | 34.1% |
Grok 4.6 High also scores 56.4% on APEX-SWE, compared with 53.6% for Grok 4.5 High and 58.8% for Fable 5 Max. On AA-Briefcase, it records 1,577, ahead of GPT-5.6 Sol Max at 1,502 and Fable 5 Max at 1,574. Its 15.8% score on Harvey LAB (Vals) is higher than the listed results for Grok 4.5 High, GPT-5.6 Sol Max and Fable 5 Max.
The comparison notes that third-party results are based on self-reported or publicly available scores. That qualification makes direct comparisons difficult, particularly where models are tested with different effort settings. One commenter also noted that the chart compares Grok 4.6 High with Max settings for the competing models.
Source: SpaceXAI on X