Model Release Radar
Price/capability frontier position, benchmark scores, and community signal
Model Release Radar
Claude Opus 4.6 (Non reasoning,
anthropicยทclosed weightsยทProprietaryยทreleased Thursday, Feb 5, 2026
No matched provider pricing is available yet.
Not enough paired price and capability data exists yet to place this model on the price/capability frontier.
This model has not been measured on DeepSWE (or any other benchmark with a real measured per-task cost) yet, so no cost/capability chart is drawn. Its Artificial Analysis benchmark scores are listed below.
Measured cost per task (DeepSWE)
This model has not been measured on DeepSWE (or any other benchmark with a real measured per-task cost) yet.
Artificial Analysis scores
Artificial Analysis scores describe capability. Compare them against estimated token spend with your cache mix; these estimates are not measured task costs.
| Benchmark | Score |
|---|---|
| AA intelligence index | 26.4 |
| GPQA Diamond (science) | 84.0% |
| Humanity's Last Exam | 19.1% |
| IFBench (instruction following) | 44.6% |
| LCR | 67.0% |
| Tau2 (tool use) | 84.8% |
| Terminal-Bench (hard) | 48.5% |
This page covers 2 reasoning-effort variants of the same model.
| Variant | AA coding index | AA intelligence index | Blended /1M (3:1 input/output) |
|---|---|---|---|
| Standard (shown above) | - | 26.4 | undisclosed |
| High effort | - | - | $10/1M |
| Variant | Coding Elo | Overall Elo | Votes |
|---|---|---|---|
| High effort | 1,535 | 1,504 | 77,193 |
| Standard | 1,535 | 1,498 | 81,769 |