Story
arxiv_cs_ai ยท Jul 24, 2026 ยท paper
arxiv.orgJul 24, 2026
original source linked
In brief
Adding procedural skills to an LLM agent is typically evaluated by average improvement in task success. However, this metric hides an important cost: skills can also make agents worse. We measure both sides by compari...
Feed lens
agentharnessevaluation