Model Release Radar
Price/capability frontier position, benchmark scores, and community signal
Model Release Radar
GLM-5.2
zai·open weights·MIT·released Tuesday, Jun 16, 2026
Frontier position
- Behind frontier DeepSWE pass@1 (agentic coding) - Muse Spark 1.2, GPT-5.6 Sol, Grok 4.6, Claude Opus 5, Grok 4.5 are priced the same or lower and at least as capable.
Cost vs. capability
Loading chart…
Benchmark scores
Measured cost per task (DeepSWE)
| Benchmark | Pass@1 | Measured cost / task | Frontier |
|---|---|---|---|
| DeepSWE (agentic coding) (max effort, 95% CI 42.0%-45.5%, 4 runs) | 43.8% | $3.92 (median $3.46) | behind |
Artificial Analysis scores (no matched cost)
Artificial Analysis scores below have no measured per-task cost, so no frontier claim is made for them here.
| Benchmark | Score |
|---|---|
| AA intelligence index | 52.6 |
| AA coding index | 68.8 |
| GPQA Diamond (science) | 89.5% |
| Humanity's Last Exam | 41.1% |
| IFBench (instruction following) | 73.3% |
| LCR | 76.7% |
| SciCode (coding) | 50.5% |
| Tau2 (tool use) | 99.1% |
| Tau2 (banking) | 34.6% |
| Terminal-Bench (hard) | 50.8% |
| Terminal-Bench 2.1 (agentic) | 77.9% |
Variants
This page covers 2 reasoning-effort variants of the same model.
| Variant | AA coding index | AA intelligence index | Price /1M |
|---|---|---|---|
| Max effort (shown above) | 68.8 | 52.6 | $2.15/1M |
| Non-reasoning | 46.5 | 34.8 | $2.15/1M |
Community signal
| Variant | Coding Elo | Overall Elo | Votes |
|---|---|---|---|
| Max effort | 1,478 | 1,467 | 32,495 |