Model Release Radar
Price/capability frontier position, benchmark scores, and community signal
Model Release Radar
GPT-5.1
openaiยทclosed weightsยทProprietaryยทreleased Thursday, Nov 13, 2025
Not enough paired price and capability data exists yet to place this model on the price/capability frontier.
This model has not been measured on DeepSWE (or any other benchmark with a real measured per-task cost) yet, so no cost/capability chart is drawn. Its Artificial Analysis benchmark scores are listed below.
Measured cost per task (DeepSWE)
This model has not been measured on DeepSWE (or any other benchmark with a real measured per-task cost) yet.
Artificial Analysis scores (no matched cost)
Artificial Analysis scores below have no measured per-task cost, so no frontier claim is made for them here.
| Benchmark | Score |
|---|---|
| AA intelligence index | 37.5 |
| AA coding index | 49.4 |
| AIME 2025 (math) | 94.0% |
| GPQA Diamond (science) | 87.3% |
| Humanity's Last Exam | 28.5% |
| IFBench (instruction following) | 72.9% |
| LCR | 76.7% |
| LiveCodeBench (coding) | 86.8% |
| MMLU-Pro (knowledge) | 87.0% |
| SciCode (coding) | 43.3% |
| Tau2 (tool use) | 81.9% |
| Tau2 (banking) | 15.9% |
| Terminal-Bench (hard) | 45.5% |
| Terminal-Bench 2.1 (agentic) | 52.4% |
This page covers 2 reasoning-effort variants of the same model.
| Variant | AA coding index | AA intelligence index | Price /1M |
|---|---|---|---|
| High effort (shown above) | 49.4 | 37.5 | $3.44/1M |
| Standard | 49.4 | 37.5 | $3.44/1M |
| Variant | Coding Elo | Overall Elo | Votes |
|---|---|---|---|
| High effort | 1,453 | 1,442 | 40,297 |
| Standard | 1,437 | 1,423 | 43,001 |