Story
arxiv_cs_lg ยท Jun 18, 2026 ยท paper
arxiv.orgJun 18, 2026
original source linked
In brief
Mainstream LLM serving systems reuse prefix work mainly through paged or radix key-value (KV) caches. This is highly effective for high-throughput, high-concurrency serving, but it manages only one positional fragment...
Feed lens
agenteval