Model Release Radar
Price/capability frontier position, benchmark scores, and community signal
Model Release Radar
Gemini 3.1 Pro Preview
google·closed weights·Proprietary·released Thursday, Feb 19, 2026
Frontier position
- Behind frontier DeepSWE pass@1 (agentic coding) - Gemini 3.7 Flash, GPT-5.6 Terra, Claude Opus 5, GPT-5.6 Sol, Grok 4.6 are priced the same or lower and at least as capable.
Cost vs. capability
Loading chart…
Benchmark scores
Measured cost per task (DeepSWE)
| Benchmark | Pass@1 | Measured cost / task | Frontier |
|---|---|---|---|
| DeepSWE (agentic coding) (high effort, 95% CI 10.2%-13.2%, 4 runs) | 11.7% | $2.14 (median $1.72) | behind |
Artificial Analysis scores (no matched cost)
Artificial Analysis scores below have no measured per-task cost, so no frontier claim is made for them here.
| Benchmark | Score |
|---|---|
| AA intelligence index | 47.7 |
| AA coding index | 68.8 |
| GPQA Diamond (science) | 94.1% |
| Humanity's Last Exam | 47.0% |
| IFBench (instruction following) | 77.1% |
| LCR | 79.0% |
| SciCode (coding) | 58.9% |
| Tau2 (tool use) | 95.6% |
| Tau2 (banking) | 21.4% |
| Terminal-Bench (hard) | 53.8% |
| Terminal-Bench 2.1 (agentic) | 73.8% |
Community signal
| Variant | Coding Elo | Overall Elo | Votes |
|---|---|---|---|
| Rating | 1,483 | 1,480 | 101,163 |