Story
arxiv_cs_cl ยท Sep 11, 2026 ยท paper
Source brief
Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models
arxiv.orgSep 11, 2026
original source linked
In brief
Large language models (LLMs) have achieved strong performance on a wide range of natural language tasks, and recent benchmarks suggest that they are increasingly adept at multi-hop reasoning. However, these benchmarks...
Feed lens
agentharnesseval