Story
arxiv_cs_ai ยท Sep 3, 2026 ยท paper
arxiv.orgSep 3, 2026
original source linked
In brief
Repository-level software engineering benchmarks have significantly advanced the evaluation of coding agents, but existing benchmarks primarily measure whether generated patches pass functional tests and overlook revi...
Feed lens
agentevaluation