Model Release Radar
Price/capability frontier position, benchmark scores, and community signal
Model Release Radar
Gemini 3.5 Flash
google·closed weights·Proprietary·released Tuesday, May 19, 2026
Frontier position
- Behind frontier DeepSWE pass@1 (agentic coding) - Claude Opus 5, GPT-5.6 Sol, Grok 4.5, Muse Spark 1.1, Gemini 3.6 Flash are priced the same or lower and at least as capable.
Cost vs. capability
Loading chart…
Benchmark scores
Measured cost per task (DeepSWE)
| Benchmark | Pass@1 | Measured cost / task | Frontier |
|---|---|---|---|
| DeepSWE (agentic coding) (high effort, 95% CI 32.1%-40.0%, 4 runs) | 36.1% | $3.45 (median $2.86) | behind |
Artificial Analysis scores (no matched cost)
Artificial Analysis scores below have no measured per-task cost, so no frontier claim is made for them here.
| Benchmark | Score |
|---|---|
| AA intelligence index | 52.0 |
| AA coding index | 70.1 |
| GPQA Diamond (science) | 92.2% |
| Humanity's Last Exam | 42.7% |
| IFBench (instruction following) | 76.3% |
| LCR | 81.0% |
| SciCode (coding) | 53.1% |
| Tau2 (tool use) | 95.3% |
| Tau2 (banking) | 32.2% |
| Terminal-Bench (hard) | 40.9% |
| Terminal-Bench 2.1 (agentic) | 78.7% |
Variants
This page covers 2 reasoning-effort variants of the same model.
| Variant | AA coding index | AA intelligence index | Price /1M |
|---|---|---|---|
| High effort (shown above) | 70.1 | 52.0 | $3.38/1M |
| Medium effort | - | 46.7 | $3.38/1M |
Community signal
| Variant | Coding Elo | Overall Elo | Votes |
|---|---|---|---|
| High effort | 1,493 | 1,483 | 32,174 |
| Medium effort | 1,484 | 1,475 | 30,593 |