LLM Digest
Subscribe

Model Release Radar

Price/capability frontier position, benchmark scores, and community signal

View as JSON

Model Release Radar

GPT-5.5

openai·closed weights·Proprietary·released Thursday, Apr 23, 2026

Frontier position
Cost vs. capability

Loading chart…

Benchmark scores

Measured cost per task (DeepSWE)

BenchmarkPass@1Measured cost / taskFrontier
DeepSWE (agentic coding) (high effort, 95% CI 61.3%-67.5%, 4 runs)64.4%$5.10 (median $4.58)behind

Artificial Analysis scores (no matched cost)

Artificial Analysis scores below have no measured per-task cost, so no frontier claim is made for them here.

BenchmarkScore
AA intelligence index54.7
AA coding index71.6
GPQA Diamond (science)93.2%
Humanity's Last Exam45.0%
IFBench (instruction following)71.6%
LCR79.0%
SciCode (coding)55.9%
Tau2 (tool use)93.0%
Tau2 (banking)36.7%
Terminal-Bench (hard)59.8%
Terminal-Bench 2.1 (agentic)79.4%
Variants

This page covers 2 reasoning-effort variants of the same model.

VariantAA coding indexAA intelligence indexPrice /1M
High effort (shown above)71.654.7$11.25/1M
Standard74.956.3$11.25/1M
Community signal
VariantCoding EloOverall EloVotes
High effort1,4951,47161,604
Standard1,4811,46662,975