LLM Digest
Subscribe

Model Release Radar

Price/capability frontier position, benchmark scores, and community signal

View as JSON

Model Release Radar

GPT-5.6 Sol

openai·closed weights·Proprietary·released Thursday, Jul 9, 2026

Frontier position
  • On frontier DeepSWE pass@1 (agentic coding) - no tracked model currently beats it on both price and capability.
Cost vs. capability

Loading chart…

Benchmark scores

Measured cost per task (DeepSWE)

BenchmarkPass@1Measured cost / taskFrontier
DeepSWE (agentic coding) (xhigh effort, 95% CI 69.9%-71.5%, 4 runs)70.7%$3.60 (median $3.15)frontier

Artificial Analysis scores (no matched cost)

Artificial Analysis scores below have no measured per-task cost, so no frontier claim is made for them here.

BenchmarkScore
AA intelligence index59.0
AA coding index78.3
GPQA Diamond (science)93.1%
Humanity's Last Exam47.3%
IFBench (instruction following)71.0%
LCR76.3%
SciCode (coding)56.0%
Tau2 (tool use)84.8%
Tau2 (banking)38.1%
Terminal-Bench (hard)61.4%
Terminal-Bench 2.1 (agentic)89.5%
Variants

This page covers 6 reasoning-effort variants of the same model.

VariantAA coding indexAA intelligence indexPrice /1M
Extra-high effort (shown above)78.359.0$8.00/1M
High effort77.257.3$8.00/1M
Low effort69.750.7$8.00/1M
Max effort77.460.9$8.00/1M
Medium effort76.355.6$8.00/1M
Non-reasoning65.141.9$8.00/1M
Community signal
VariantCoding EloOverall EloVotes
Extra-high effort1,4901,45421,537