Story
arxiv_cs_ai ยท Sep 16, 2026 ยท paper
arxiv.orgSep 16, 2026
original source linked
In brief
Reliable agent evaluation is complicated by automatic harness optimization, which repeatedly uses a released benchmark $B_{\mathrm{rel}}$ to guide a Proposer that edits prompts, memory, retrieval, tools, and control c...
Feed lens
agentharnessevaluation