What does PagedAttention change in LLM KV allocation?

PagedAttention lets a request’s attention cache occupy separate fixed-size memory blocks instead of one large contiguous allocation. A block table maps the request’s logical token positions to physical blocks. This reduces the need to reserve a request’s maximum possible sequence length in advance.

The main benefit is more flexible KV allocation for concurrent requests. It does not reduce the key/value payload required for each stored token.

Calculate unused block space

Consider an illustrative implementation with 16-token blocks. A 35-token sequence needs three blocks, providing space for 48 tokens. Thirteen positions in the last block are unused: about 27% of the allocated positions for this short sequence.

Only the last block needs partial filling. As sequences grow, that unused space becomes a smaller fraction of their allocation. The actual block size and supported layouts depend on the runtime, attention backend, and version.

The PagedAttention paper describes the mapping and allocation design. It allocates additional blocks as a sequence grows and releases them when the request finishes, rather than reserving a large contiguous range per request.

Sharing requires more than allocation

Multiple requests can reference the same compatible cache block. Prefix caching finds reusable blocks for identical prefixes; copy-on-write handling lets requests diverge without corrupting shared state. A paged layout enables this, but does not automatically provide a cache lookup policy or guarantee reuse across requests.

The original vLLM introduction reported 60–80% waste in the earlier compared allocation approaches and below 4% in its measured PagedAttention workloads. The short-sequence calculation above shows why below 4% is not a universal upper bound. Its large throughput comparisons also include full-runtime differences, so they should not be attributed to allocation alone.

Measure allocated blocks, used token positions, evictions or preemptions, and concurrent request capacity. Compare the same lengths and cache precision. If raw KV data already exceeds memory, block allocation cannot make that payload disappear; fewer KV heads, supported cache quantization, or a smaller workload may also be required.

Engineering Guide: PagedAttention includes a block-table example.