LLM Digest
Subscribe

Model Release Radar

Price/capability frontier position, benchmark scores, and community signal

View as JSON

Model Release Radar

Gemini 3.8 Flash

google·released Wednesday, Sep 2, 2026

Frontier position
  • On frontier DeepSWE pass@1 (agentic coding) - no tracked model currently beats it on both price and capability.
Cost vs. capability

Loading chart…

Benchmark scores

Measured cost per task (DeepSWE)

BenchmarkPass@1Measured cost / taskFrontier
DeepSWE (agentic coding) (high effort, 95% CI 72.4%-75.2%, 4 runs)73.8%$2.36 (median $2.11)frontier

Artificial Analysis scores (no matched cost)

Artificial Analysis scores below have no measured per-task cost, so no frontier claim is made for them here.

BenchmarkScore
AA intelligence index58.7
AA coding index76.3
GPQA Diamond (science)95.3%
Humanity's Last Exam47.8%
LCR81.0%
SciCode (coding)53.6%
Tau2 (banking)44.9%
Terminal-Bench 2.1 (agentic)87.6%
Variants

This page covers 3 reasoning-effort variants of the same model.

VariantAA coding indexAA intelligence indexPrice /1M
High effort (shown above)76.358.7$1.50/1M
Low effort73.551.7$1.50/1M
Medium effort74.156.6$1.50/1M
Community signal

No LMArena community rating is tracked for this model yet.