Model Release Radar
Price/capability frontier position, benchmark scores, and community signal
Model Release Radar
GPT-5.4
openai·closed weights·Proprietary·released Thursday, Mar 5, 2026
Frontier position
- Behind frontier DeepSWE pass@1 (agentic coding) - Grok 4.6, GPT-5.5, Kimi K3, Claude Opus 4.8, GLM-5.3 are priced the same or lower and at least as capable.
Cost vs. capability
Loading chart…
Benchmark scores
Measured cost per task (DeepSWE)
| Benchmark | Pass@1 | Measured cost / task | Frontier |
|---|---|---|---|
| DeepSWE (agentic coding) (xhigh effort, 95% CI 50.3%-53.3%, 4 runs) | 51.8% | $5.65 (median $4.35) | behind |
Artificial Analysis scores (no matched cost)
Artificial Analysis scores below have no measured per-task cost, so no frontier claim is made for them here.
| Benchmark | Score |
|---|---|
| AA intelligence index | 53.1 |
| AA coding index | 71.1 |
| GPQA Diamond (science) | 92.0% |
| Humanity's Last Exam | 43.7% |
| IFBench (instruction following) | 73.9% |
| LCR | 77.7% |
| SciCode (coding) | 56.6% |
| Tau2 (tool use) | 87.1% |
| Tau2 (banking) | 39.6% |
| Terminal-Bench (hard) | 57.6% |
| Terminal-Bench 2.1 (agentic) | 78.3% |
Variants
This page covers 2 reasoning-effort variants of the same model.
| Variant | AA coding index | AA intelligence index | Price /1M |
|---|---|---|---|
| Standard (shown above) | 71.1 | 53.1 | $5.62/1M |
| High effort | - | - | undisclosed |
Community signal
| Variant | Coding Elo | Overall Elo | Votes |
|---|---|---|---|
| High effort | 1,496 | 1,470 | 60,614 |
| Standard | 1,480 | 1,453 | 63,594 |