Model Release Radar
Price/capability frontier position, benchmark scores, and community signal
Model Release Radar
Grok 4.6
xai·closed weights·Proprietary·released Wednesday, Aug 12, 2026
Frontier position
- Behind frontier DeepSWE pass@1 (agentic coding) - GPT-5.6 Terra, Claude Opus 5, GPT-5.6 Sol, GLM-5.3, GPT-5.6 Luna are priced the same or lower and at least as capable.
Cost vs. capability
Loading chart…
Benchmark scores
Measured cost per task (DeepSWE)
| Benchmark | Pass@1 | Measured cost / task | Frontier |
|---|---|---|---|
| DeepSWE (agentic coding) (high effort, 95% CI 63.6%-66.7%, 4 runs) | 65.2% | $4.38 (median $3.69) | behind |
Artificial Analysis scores (no matched cost)
Artificial Analysis scores below have no measured per-task cost, so no frontier claim is made for them here.
| Benchmark | Score |
|---|---|
| AA intelligence index | 60.9 |
| AA coding index | 76.8 |
| GPQA Diamond (science) | 94.9% |
| Humanity's Last Exam | 42.9% |
| LCR | 75.0% |
| SciCode (coding) | 53.6% |
| Tau2 (banking) | 50.7% |
| Terminal-Bench 2.1 (agentic) | 88.4% |
Variants
This page covers 4 reasoning-effort variants of the same model.
| Variant | AA coding index | AA intelligence index | Price /1M |
|---|---|---|---|
| High effort (shown above) | 76.8 | 60.9 | $3.00/1M |
| Extra-high effort | 75.9 | 60.0 | $3.00/1M |
| Low effort | 66.3 | 51.7 | $3.00/1M |
| Medium effort | 74.4 | 59.0 | $3.00/1M |
Community signal
| Variant | Coding Elo | Overall Elo | Votes |
|---|---|---|---|
| High effort | 1,476 | 1,444 | 3,471 |