Story
arxiv_cs_cl ยท Jul 29, 2026 ยท paper
Source brief
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
arxiv.orgJul 29, 2026
original source linked
In brief
Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows...
Feed lens
agentevaluation