ValueDiff: Value-Geometric KV Cache Eviction for Sink-Suppressed LLMs
Modern LLMs with QK-normalization, gated attention, learned attention sinks, or logit softcapping exhibit weaker persistent attention sinks, on which existing KV cache eviction methods primarily rely. We observe that across these models, weaker sinks co-occur with greater value-vector dispersion relative to key-vector dispersion. Motivated by this value-side dispersion, we present ValueDiff, a…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.