Story
arxiv_cs_cl ยท Aug 4, 2026 ยท paper
arxiv.orgAug 4, 2026
original source linked
In brief
Large language models (LLMs), and the agents built on top of them, are now benchmarked heavily on whether they can finish a task -- fix a bug, drive a browser, operate a GUI. A complementary social ability, namely how...
Feed lens
agenteval