How much KV cache memory does an LLM request require?
For a conventional full-attention transformer, KV-cache memory grows with the number of cached tokens, layers, key/value heads, and bytes per value. Multiply the per-sequence allocation by concurrent sequences to estimate the raw payload. Then reserve additional memory for weights and runtime work.
This estimate helps serving engineers size a workload. Multi-head latent attention (MLA), sliding-window, hybrid, and recurrent caches need architecture-specific estimates.
Use KV heads rather than query heads
The cache stores earlier tokens’ key and value projections so the model does not recompute them at every generation step. Dense attention still reads a growing token history. For full-attention caches with separate key/value heads, as in MHA, MQA, or GQA, and identical-length sequences:
cache bytes = 2 × layers × KV_heads × head_dimension
× cached_tokens × concurrent_sequences × bytes_per_element
The factor two counts keys and values. For unequal sequences, use the sum of their cached token counts. Read an explicit head_dim from the model configuration when available; hidden_size / num_attention_heads is only a fallback for architectures that use that relationship.
Meta’s Llama 3 architecture table provides the dimensions for these examples, using a two-byte FP16/BF16 cache:
| Model | Layers / KV heads / head dimension | Tokens per sequence | Raw cache per sequence |
|---|---|---|---|
| Llama 3.1 8B | 32 / 8 / 128 | 8,192 | 1 GiB |
| Llama 3.1 8B | 32 / 8 / 128 | 131,072 | 16 GiB |
| Llama 3.1 70B | 80 / 8 / 128 | 4,096 | 1.25 GiB |
| Llama 3.1 70B | 80 / 8 / 128 | 131,072 | 40 GiB |
A 40 GiB cache budget holds at most 32 full 4,096-token 70B sequences by this arithmetic. This counts cache payload only; choose the concurrency limit after reserving runtime memory.
Check the actual server allocation
Weights, temporary buffers, activations, partially filled blocks, and reserved capacity require more memory. Distributed servers may shard or replicate KV heads. A supported one-byte cache halves this raw payload, but scales, implementation support, and quality can alter usable capacity.
Measure peak allocation and preemption with representative prompt and output lengths. PagedAttention improves block allocation; it does not eliminate the stored key/value data or make sequence length irrelevant.
Engineering Guide: KV-cache memory includes an executable capacity calculation.