Explainer
What Is the KV Cache in LLMs? Why Context Length Costs Memory
Last updated: October 2, 2026 — model data is refreshed automatically from the id8 dataset
You load a model that should fit in your GPU, raise the context length, and it runs out of memory. The cause is the KV cache: memory the model uses for every token in the conversation, on top of its weights. This page explains what it is and how to budget for it.
Set any context length and see weights and cache separately.
See KV cache memory for your model →The problem it solves
A language model writes one token at a time. To choose the next token, each attention layer compares the current token with every earlier token in the context.
Done naively, the model would redo all the work for all earlier tokens at every step. Writing token 1,000 would repeat the work for tokens 1 to 999.
The KV cache removes that repetition. When a token is first processed, each layer computes two vectors for it, a key and a value, and stores them. For every later token, the model reads the stored vectors and computes only the new token's part. Generation becomes far faster, and the price is memory.
Why it grows with context
One key and one value are stored per token, per layer, per attention head. So the cache size is roughly:
context length × number of layers × key-value heads × head size × 2 × bytes per number
Everything in that product except the context length is fixed by the model. So for a given model, the cache is proportional to the context length: twice the context, twice the cache.
What that looks like in practice
Total memory for Qwen 3 14B at different quantizations and context lengths:
| Quantization | VRAM at 4K context | VRAM at 8K context | VRAM at 32K context |
|---|---|---|---|
| Q3_K_M | 8.4 GB | 9.0 GB | 12.8 GB |
| Q4_K_M | 10.0 GB | 10.6 GB | 14.4 GB |
| Q5_K_M | 11.4 GB | 12.1 GB | 15.8 GB |
| Q6_K | 12.9 GB | 13.5 GB | 17.3 GB |
| Q8_0 | 16.2 GB | 16.9 GB | 20.6 GB |
Read down a column: quantization changes the weights. Read across a row: the increase is the KV cache. At long contexts the cache can rival the weights themselves.
A larger model for comparison, Qwen 3 32B:
| Quantization | VRAM at 4K context | VRAM at 8K context | VRAM at 32K context |
|---|---|---|---|
| Q3_K_M | 16.7 GB | 17.7 GB | 23.7 GB |
| Q4_K_M | 20.2 GB | 21.2 GB | 27.2 GB |
| Q5_K_M | 23.2 GB | 24.2 GB | 30.2 GB |
| Q6_K | 26.5 GB | 27.5 GB | 33.5 GB |
| Q8_0 | 33.7 GB | 34.7 GB | 40.7 GB |
What changes the size of the cache
- Context length. The main factor, and the one you control.
- Model architecture. More layers and wider attention mean a bigger cache per token. Many current models use grouped-query attention, which shares keys and values across groups of heads and cuts the cache several times over.
- Sliding-window attention. Some models limit certain layers to a recent window of tokens, which caps the cache for those layers.
- Cache precision. The cache is usually stored in 16-bit numbers. Some runtimes can store it in 8-bit, halving it with a small quality cost.
- Batching. Serving several conversations at once needs a cache for each.
Weight quantization, such as Q4 against Q8, does not change the cache.
Practical rules
- Set the context to what you need. If your chats are short, an 8K context is plenty. Do not set 128K because the model supports it; the memory is reserved when the model loads.
- Budget weights plus cache. A model whose weights just fit will not run with a long context.
- For agents and document work, where context must be long, choose a smaller model or a lower quantization to make room.
- Prefer models with efficient attention if long context is your main use.
- Check before downloading. Can I Run LLM? shows the total for any model, quantization and context on your GPU.
The KV cache and prompt caching are different things
- The KV cache lives in your GPU's memory during one generation. It is what this page describes.
- Prompt caching is a feature of API providers. They keep the processed form of a long, repeated prompt prefix on their servers for a short time and bill repeated input at a discount.
They rest on the same idea, not recomputing what has been computed, at different scales.
In fine-tuning
Training does not use a KV cache in the same way. It stores activations for every token so gradients can be computed, which is a larger cost with the same shape: proportional to sequence length. See How Much VRAM Do You Need to Fine-Tune an LLM?.
Frequently asked questions
What does KV stand for?
Key and value. They are two sets of numbers each attention layer computes for every token, and the cache stores them so they are not recomputed.
How much memory does the KV cache use?
It grows in proportion to context length. For Qwen 3 14B at Q4_K_M, total memory is about 10.0 GB with a 4K context and about 14.4 GB with a 32K context; the difference is the cache.
Does quantizing the model shrink the KV cache?
No. Quantization shrinks the weights. The cache is a separate block whose size depends on the model's attention layers and the context length. Some tools can store the cache itself at lower precision as a separate option.
Is memory used for the context I set or the context I use?
Local tools usually reserve the cache for the full context length you set when the model loads, whether or not you fill it.
Why is the first reply slow after a long prompt?
The model has to process every token of the prompt once to fill the cache. After that, each new token is fast.