How prefix caching reuses LLM prompts and cached tokens
Prefix caching reuses attention state from an earlier compatible prompt prefix. A new request processes only the remaining uncached portion instead of repeating all prompt computation. It helps workloads with repeated system instructions, examples, documents, or conversation prefixes.
The match is over tokens and model state. Similar wording, repeated text later in a prompt, or the same tokens under a different adapter do not establish a valid hit.
What must match
In a causal transformer, a token’s cached keys and values depend on the tokens before it. Reusing a later repeated passage after a changed earlier passage would reuse different context-dependent states.
| Candidate reuse | Check |
|---|---|
| Identical system prompt | Exact token IDs and compatible model state |
| Repeated document | Identical preceding prefix as well as the document |
| Continued conversation | Unchanged earlier messages, template, and tokens |
| Same text with another adapter | Adapter identity and the cache’s compatibility rules |
| Multimodal prompt | Image/media identity, not placeholder tokens alone |
vLLM’s prefix-cache design hashes a block’s tokens, its parent prefix, and extra identifiers such as LoRA and multimodal state. It reuses complete blocks; an incomplete final block is not automatically a cache hit. Cache salts can also separate requests that should not share cached state.
SGLang’s RadixAttention organizes reusable prefixes in a radix tree. The representation differs, but compatible prefix state is still required.
Measure avoided prefill work
Prefix caching mainly avoids repeated prompt computation. It does not remove decoding or its attention over the available history. Saved computation also does not imply that every request’s TTFT improves: queueing and cold-cache requests remain.
Compare cold requests with warmed repeated-prefix requests separately. Record reused tokens, not just the fraction of requests with any hit. A hit covering 32 tokens and one covering 8,000 tokens avoid different amounts of work. Measure retained cache memory, evictions, and latency by prefix length.
Put stable instructions before request-specific data when that order preserves the intended prompt. Keep tenant and adapter compatibility in the cache policy. Report the cache condition whenever you publish serving results.
Engineering Guide: prefix caching relates the reuse mechanism to KV allocation.