LLM Digest
Subscribe

Model Release Radar

Price/capability frontier position, benchmark scores, and community signal

View as JSON

Model Release Radar

Gemini 3.5 Flash

google·closed weights·Proprietary·released Tuesday, May 19, 2026

Frontier position
Cost vs. capability

Loading chart…

Benchmark scores

Measured cost per task (DeepSWE)

BenchmarkPass@1Measured cost / taskFrontier
DeepSWE (agentic coding) (high effort, 95% CI 32.1%-40.0%, 4 runs)36.1%$3.45 (median $2.86)behind

Artificial Analysis scores (no matched cost)

Artificial Analysis scores below have no measured per-task cost, so no frontier claim is made for them here.

BenchmarkScore
AA intelligence index52.0
AA coding index70.1
GPQA Diamond (science)92.2%
Humanity's Last Exam42.7%
IFBench (instruction following)76.3%
LCR81.0%
SciCode (coding)53.1%
Tau2 (tool use)95.3%
Tau2 (banking)32.2%
Terminal-Bench (hard)40.9%
Terminal-Bench 2.1 (agentic)78.7%
Variants

This page covers 2 reasoning-effort variants of the same model.

VariantAA coding indexAA intelligence indexPrice /1M
High effort (shown above)70.152.0$3.38/1M
Medium effort-46.7$3.38/1M
Community signal
VariantCoding EloOverall EloVotes
High effort1,4931,48332,174
Medium effort1,4841,47530,593