Model Release Radar
Price/capability frontier position, benchmark scores, and community signal
Model Release Radar
Claude Sonnet 4.6
anthropic·closed weights·Proprietary·released Tuesday, Feb 17, 2026
Frontier position
- Behind frontier DeepSWE pass@1 (agentic coding) - Grok 4.6, GPT-5.5, Kimi K3, Claude Opus 4.8, Claude Sonnet 5 are priced the same or lower and at least as capable.
Cost vs. capability
Loading chart…
Benchmark scores
Measured cost per task (DeepSWE)
| Benchmark | Pass@1 | Measured cost / task | Frontier |
|---|---|---|---|
| DeepSWE (agentic coding) (high effort, 95% CI 25.8%-34.0%, 4 runs) | 29.9% | $5.52 (median $4.87) | behind |
Artificial Analysis scores (no matched cost)
Artificial Analysis scores below have no measured per-task cost, so no frontier claim is made for them here.
| Benchmark | Score |
|---|---|
| AA intelligence index | 36.8 |
| GPQA Diamond (science) | 79.9% |
| Humanity's Last Exam | 13.3% |
| IFBench (instruction following) | 41.2% |
| LCR | 62.3% |
| SciCode (coding) | 46.9% |
| Tau2 (tool use) | 79.5% |
| Terminal-Bench (hard) | 46.2% |
Community signal
| Variant | Coding Elo | Overall Elo | Votes |
|---|---|---|---|
| Rating | 1,504 | 1,458 | 66,354 |