Model Release Radar
Price/capability frontier position, benchmark scores, and community signal
Model Release Radar
GPT-5.6 Sol
openai·closed weights·Proprietary·released Thursday, Jul 9, 2026
Frontier position
- On frontier DeepSWE pass@1 (agentic coding) - no tracked model currently beats it on both price and capability.
Cost vs. capability
Loading chart…
Benchmark scores
Measured cost per task (DeepSWE)
| Benchmark | Pass@1 | Measured cost / task | Frontier |
|---|---|---|---|
| DeepSWE (agentic coding) (xhigh effort, 95% CI 69.9%-71.5%, 4 runs) | 70.7% | $3.60 (median $3.15) | frontier |
Artificial Analysis scores (no matched cost)
Artificial Analysis scores below have no measured per-task cost, so no frontier claim is made for them here.
| Benchmark | Score |
|---|---|
| AA intelligence index | 59.0 |
| AA coding index | 78.3 |
| GPQA Diamond (science) | 93.1% |
| Humanity's Last Exam | 47.3% |
| IFBench (instruction following) | 71.0% |
| LCR | 76.3% |
| SciCode (coding) | 56.0% |
| Tau2 (tool use) | 84.8% |
| Tau2 (banking) | 38.1% |
| Terminal-Bench (hard) | 61.4% |
| Terminal-Bench 2.1 (agentic) | 89.5% |
Variants
This page covers 6 reasoning-effort variants of the same model.
| Variant | AA coding index | AA intelligence index | Price /1M |
|---|---|---|---|
| Extra-high effort (shown above) | 78.3 | 59.0 | $8.00/1M |
| High effort | 77.2 | 57.3 | $8.00/1M |
| Low effort | 69.7 | 50.7 | $8.00/1M |
| Max effort | 77.4 | 60.9 | $8.00/1M |
| Medium effort | 76.3 | 55.6 | $8.00/1M |
| Non-reasoning | 65.1 | 41.9 | $8.00/1M |
Community signal
| Variant | Coding Elo | Overall Elo | Votes |
|---|---|---|---|
| Extra-high effort | 1,490 | 1,454 | 21,537 |