LLM Digest
Subscribe

Model Release Radar

Price/capability frontier position, benchmark scores, and community signal

View as JSON

Model Release Radar

Kimi K2.7 Code

moonshot·released Friday, Jun 12, 2026

Frontier position
Cost vs. capability

Loading chart…

Benchmark scores

Measured cost per task (DeepSWE)

BenchmarkPass@1Measured cost / taskFrontier
DeepSWE (agentic coding) (95% CI 30.0%-31.0%, 4 runs)30.5%$2.82 (median $2.20)behind

Artificial Analysis scores (no matched cost)

Artificial Analysis scores below have no measured per-task cost, so no frontier claim is made for them here.

BenchmarkScore
AA intelligence index43.0
AA coding index60.8
GPQA Diamond (science)89.6%
Humanity's Last Exam35.0%
IFBench (instruction following)63.1%
LCR75.0%
SciCode (coding)47.5%
Tau2 (tool use)90.1%
Tau2 (banking)20.2%
Terminal-Bench (hard)44.7%
Terminal-Bench 2.1 (agentic)67.4%
Community signal

No LMArena community rating is tracked for this model yet.