LLM Digest
Subscribe

Model Release Radar

Price/capability frontier position, benchmark scores, and community signal

View as JSON

Model Release Radar

GPT-5.6 Terra

openai·closed weights·Proprietary·released Thursday, Jul 9, 2026

Frontier position
Cost vs. capability

Loading chart…

Benchmark scores

Measured cost per task (DeepSWE)

BenchmarkPass@1Measured cost / taskFrontier
DeepSWE (agentic coding) (xhigh effort, 95% CI 58.1%-62.3%, 4 runs)60.2%$1.70 (median $1.49)behind

Artificial Analysis scores (no matched cost)

Artificial Analysis scores below have no measured per-task cost, so no frontier claim is made for them here.

BenchmarkScore
AA intelligence index52.8
AA coding index70.6
GPQA Diamond (science)90.8%
Humanity's Last Exam41.9%
IFBench (instruction following)66.3%
LCR75.0%
SciCode (coding)51.6%
Tau2 (tool use)80.4%
Tau2 (banking)29.7%
Terminal-Bench (hard)62.9%
Terminal-Bench 2.1 (agentic)80.1%
Variants

This page covers 6 reasoning-effort variants of the same model.

VariantAA coding indexAA intelligence indexPrice /1M
Extra-high effort (shown above)70.652.8$4.50/1M
High effort67.150.1$4.50/1M
Low effort58.141.3$4.50/1M
Max effort76.756.6$4.50/1M
Medium effort64.746.8$4.50/1M
Non-reasoning52.334.6$4.50/1M
Community signal
VariantCoding EloOverall EloVotes
Extra-high effort1,4871,44722,345