Thinking Outside the Box: Retention and Transmission of Information in Sliding-Window KV Inference
Sliding-window KV inference refers to processing a sequence incrementally while retaining only a fixed-size cache of recent key and value states. It can be applied to pretrained causal transformers at inference time without additional training, while its KV-cache memory remains fixed as more tokens are processed. Because cached states are computed in the context of earlier tokens, they may carry…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.