LLM Digest
Subscribe

Model Release Radar

Price/capability frontier position, benchmark scores, and community signal

View as JSON

Model Release Radar

GLM-5.2

zai·open weights·MIT·released Tuesday, Jun 16, 2026

Frontier position
Cost vs. capability

Loading chart…

Benchmark scores

Measured cost per task (DeepSWE)

BenchmarkPass@1Measured cost / taskFrontier
DeepSWE (agentic coding) (max effort, 95% CI 42.0%-45.5%, 4 runs)43.8%$3.92 (median $3.46)behind

Artificial Analysis scores (no matched cost)

Artificial Analysis scores below have no measured per-task cost, so no frontier claim is made for them here.

BenchmarkScore
AA intelligence index52.6
AA coding index68.8
GPQA Diamond (science)89.5%
Humanity's Last Exam41.1%
IFBench (instruction following)73.3%
LCR76.7%
SciCode (coding)50.5%
Tau2 (tool use)99.1%
Tau2 (banking)34.6%
Terminal-Bench (hard)50.8%
Terminal-Bench 2.1 (agentic)77.9%
Variants

This page covers 2 reasoning-effort variants of the same model.

VariantAA coding indexAA intelligence indexPrice /1M
Max effort (shown above)68.852.6$2.15/1M
Non-reasoning46.534.8$2.15/1M
Community signal
VariantCoding EloOverall EloVotes
Max effort1,4781,46732,495