Story

arxiv_cs_ai ยท May 12, 2026 ยท paper

Source brief

KV-Fold: One-Step KV-Cache Recurrence for Long-Context Inference

arxiv.orgMay 12, 2026
original source linked

In brief

We introduce KV-Fold, a simple, training-free long-context inference protocol that treats the key-value (KV) cache as the accumulator in a left fold over sequence chunks. At each step, the model processes the next chu...

Continue reading

Read the original at arxiv.org โ†’Open in live feed

Earlier in this thread 2 items