LLM Digest
Subscribe

Model Release Radar

Price/capability frontier position, benchmark scores, and community signal

View as JSON

Model Release Radar

GLM-5.3-Flash

zai·open weights·MIT·released Wednesday, Aug 26, 2026

Frontier position
  • On frontier DeepSWE pass@1 (agentic coding) - no tracked model currently beats it on both price and capability.
Cost vs. capability

Loading chart…

Benchmark scores

Measured cost per task (DeepSWE)

BenchmarkPass@1Measured cost / taskFrontier
DeepSWE (agentic coding) (max effort)63.4%$0.241frontier

Artificial Analysis scores (no matched cost)

Artificial Analysis scores below have no measured per-task cost, so no frontier claim is made for them here.

BenchmarkScore
AA intelligence index57.5
AA coding index71.5
GPQA Diamond (science)91.2%
Humanity's Last Exam39.9%
LCR78.0%
SciCode (coding)46.1%
Tau2 (banking)47.2%
Terminal-Bench 2.1 (agentic)84.3%
Community signal
VariantCoding EloOverall EloVotes
Rating1,5121,4662,424