Story
arxiv_cs_lg ยท Oct 5, 2026 ยท paper
arxiv.orgOct 5, 2026
original source linked
In brief
Long context inference with large language models becomes increasingly expensive as attention must operate over an ever growing KV cache. Page sparse attention reduces this cost by representing each KV page compactly...
Feed lens
eval