Story

arxiv_cs_lg ยท Oct 5, 2026 ยท paper

Source brief

OVAL: Output-Aware Local Page Bases for KV Cache Retrieval

arxiv.orgOct 5, 2026
original source linked

In brief

Long context inference with large language models becomes increasingly expensive as attention must operate over an ever growing KV cache. Page sparse attention reduces this cost by representing each KV page compactly...

Feed lens
eval

Continue reading

Read the original at arxiv.org โ†’Open in live feedRead that dayโ€™s brief

Earlier in this thread 4 items