LLM Digest
Subscribe

Model Release Radar

Price/capability frontier position, benchmark scores, and community signal

View as JSON

Model Release Radar

Claude Opus 5

anthropic·closed weights·Proprietary·released Friday, Jul 24, 2026

Frontier position
  • On frontier DeepSWE pass@1 (agentic coding) - no tracked model currently beats it on both price and capability.
Cost vs. capability

Loading chart…

Benchmark scores

Measured cost per task (DeepSWE)

BenchmarkPass@1Measured cost / taskFrontier
DeepSWE (agentic coding) (max effort)73.7%$11.84frontier

Artificial Analysis scores (no matched cost)

Artificial Analysis scores below have no measured per-task cost, so no frontier claim is made for them here.

BenchmarkScore
AA intelligence index61.5
AA coding index76.5
GPQA Diamond (science)93.7%
Humanity's Last Exam52.8%
LCR76.3%
SciCode (coding)54.3%
Tau2 (banking)44.7%
Terminal-Bench 2.1 (agentic)87.6%
Variants

This page covers 5 reasoning-effort variants of the same model.

VariantAA coding indexAA intelligence indexPrice /1M
High effort (shown above)76.561.5$10/1M
Extra-high effort77.062.5$10/1M
Low effort66.952.5$10/1M
Max effort78.063.1$10/1M
Medium effort74.358.6$10/1M
Community signal
VariantCoding EloOverall EloVotes
High effort1,5311,50431,570
Max effort1,5281,50515,398