Understanding Its Purpose First
The KV cache stores the attention keys and values of processed tokens, reusing them in subsequent generation to reduce redundant computation.
Understanding with an Example
When generating text, previously processed context does not need to have all historical keys and values recomputed for every new token. The cache preserves the reusable portion, and the new step appends its own keys and values before completing the current attention computation.
Calculate historical K/V for each layer
Next step: read existing cache
Calculate new K/V and continue
Going a Step Deeper
In common causal autoregressive models, older positions do not change their historical keys and values when future new tokens appear, so they can be reused across layers. Current queries still need to interact with accessible historical content—the cache does not completely eliminate computation. The cache typically grows with the sequence length, consuming GPU memory or RAM.
What It Doesn't Mean
The KV cache is an inference computation state, not user preference memory or a database. It does not automatically expand the context that the model claims to support; different architectures, sliding windows, or caching strategies each have their own limitations.
Continue Learning
References
- Hugging Face · Caching: Layer-wise reuse of historical K/V in autoregressive inference; attention Q/K/V and cache memory overhead.
Verified on 2026-09-09; original papers are used to explain the mechanism, and the examples in the text are for instructional purposes.