Story

arxiv_cs_cl ยท Sep 11, 2026 ยท paper

Source brief

Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models

arxiv.orgSep 11, 2026
original source linked

In brief

Large language models (LLMs) have achieved strong performance on a wide range of natural language tasks, and recent benchmarks suggest that they are increasingly adept at multi-hop reasoning. However, these benchmarks...

Feed lens
agentharnesseval

Continue reading

Read the original at arxiv.org โ†’Open in live feedRead that dayโ€™s brief

Earlier in this thread 4 items