Story
arxiv_cs_cl ยท Sep 24, 2026 ยท paper
arxiv.orgSep 24, 2026
original source linked
In brief
Vision-Language Models (VLMs) are increasingly used for long-document processing, where the inputs combine text with charts, tables, figures, and complex layouts. Deploying them means choosing how to feed the document...
Feed lens
agenticeval