Story
arxiv_cs_ai ยท Jun 16, 2026 ยท paper
arxiv.orgJun 16, 2026
original source linked
In brief
As high-quality public web corpora become increasingly exhausted, clean long-context documents have become a scarce and expensive source of training data for large language models (LLMs). Existing long-context corpora...
Feed lens
evaluation