LLM Digest
Subscribe

Model Release Radar

Price/capability frontier position, benchmark scores, and community signal

View as JSON

Model Release Radar

GPT-5.4

openai·closed weights·Proprietary·released Thursday, Mar 5, 2026

Frontier position
Cost vs. capability

Loading chart…

Benchmark scores

Measured cost per task (DeepSWE)

BenchmarkPass@1Measured cost / taskFrontier
DeepSWE (agentic coding) (xhigh effort, 95% CI 50.3%-53.3%, 4 runs)51.8%$5.65 (median $4.35)behind

Artificial Analysis scores (no matched cost)

Artificial Analysis scores below have no measured per-task cost, so no frontier claim is made for them here.

BenchmarkScore
AA intelligence index53.1
AA coding index71.1
GPQA Diamond (science)92.0%
Humanity's Last Exam43.7%
IFBench (instruction following)73.9%
LCR77.7%
SciCode (coding)56.6%
Tau2 (tool use)87.1%
Tau2 (banking)39.6%
Terminal-Bench (hard)57.6%
Terminal-Bench 2.1 (agentic)78.3%
Variants

This page covers 2 reasoning-effort variants of the same model.

VariantAA coding indexAA intelligence indexPrice /1M
Standard (shown above)71.153.1$5.62/1M
High effort--undisclosed
Community signal
VariantCoding EloOverall EloVotes
High effort1,4961,47060,614
Standard1,4801,45363,594