Prefill vs decode: why LLM inference has two phases

Prefill processes a prompt and creates its attention cache. Decode uses that cache to generate subsequent tokens. Long or batched prefill can use large matrix operations efficiently; small-batch decode often spends more time reading weights and cached keys and values.

This distinction helps engineers diagnose a slow first token separately from slow generation after it.

Identify which phase is slow

Ordinary autoregressive prefill produces the logits used to sample the first output token. Later decode steps add one token at a time. Time to first token also includes queueing, tokenization, scheduling, and network delivery, so a large TTFT does not prove prefill itself is slow.

SymptomMeasure before changing the server
First token gets slower with longer promptsQueue time and prefill duration by input length
Later tokens slow down with longer historiesDecode time and KV-cache traffic
A long new prompt pauses other streamsScheduler iterations and inter-token intervals
Separate servers spend time transferring stateKV-transfer bytes, bandwidth, and queueing

Prefill reuses weights across prompt tokens. The implementation may still load matrix tiles more than once; it does not guarantee one literal HBM read of every weight. Splitwise describes how these phases use different hardware resources.

What chunked prefill changes

Chunked prefill processes a long prompt over several scheduler iterations. vLLM’s documented scheduler admits running decode requests first and uses the remaining iteration token budget for prefill. The actual chunk size follows the available budget.

This can reduce interruptions to existing streams while increasing the new request’s TTFT. Smaller chunks also add scheduling work and reread earlier cache state. Sarathi-Serve evaluates this trade-off; its approximately 25% additional prefill time at 512-token chunks was measured for Yi-34B with tensor parallelism of two.

Disaggregated serving instead places prefill and decode on separate GPU pools. It can tune each phase independently, but must transfer the request’s KV cache. Check transfer cost and pool utilization before adopting that design. Chunking is a scheduler change; disaggregation also changes deployment and communication.

Engineering Guide: prefill versus decode includes timelines for both approaches.