Why is LLM decode memory-bound and prefill compute-bound?
LLM decode is often memory-bound because a small batch reuses each model weight only a few times before loading the next weights. Long or sufficiently batched prefill reuses weights across many prompt tokens, so matrix computation can become the limiting resource. These are workload conditions, rather than permanent properties of either phase.
Arithmetic intensity helps decide whether to reduce data movement or improve matrix execution.
Compare calculations with bytes moved
Arithmetic intensity is the number of floating-point operations performed per byte transferred from memory. The roofline model compares it with peak compute divided by memory bandwidth.
NVIDIA’s H100 SXM specifications give approximately 989 TFLOPS for dense FP16/BF16 Tensor Core computation and 3.35 TB/s memory bandwidth. The listed 1,979-TFLOPS figure assumes structured sparsity. Their ratio is about 295 FLOPs per byte. This is a theoretical bound using matching hardware specifications.
A simplified dense matrix-vector multiply performs about two operations per two-byte weight: roughly one FLOP per byte. A batch-one decoder in this regime cannot use the GPU’s peak matrix throughput. Larger batches let a loaded weight contribute to several token computations. Prefill similarly reuses weights across prompt tokens. Splitwise describes these different phase characteristics.
Choose a change from the measured limit
| Observation | Change to investigate | What can prevent a gain |
|---|---|---|
| Small-batch decode spends time moving weights | Supported weight quantization; larger decode batches | Quantization kernels, quality loss, latency limits |
| Long-context decode spends time reading attention history | Fewer KV heads; supported KV quantization; efficient attention | Model architecture and cache precision support |
| Large prefill uses matrix units heavily | Efficient matrix kernels; supported lower-precision compute | Short prompts, small batches, other operations |
| Many short kernels leave gaps in GPU execution | Fusion or reduced launch overhead | Added register/shared-memory use |
Measure both first-token latency and later token intervals after each change. A larger batch can improve total tokens per second while making each request slower. The hardware ratio identifies a possible limit; a trace establishes which resource your workload actually waits for.
Engineering Guide: memory-bound versus compute-bound inference includes the roofline diagram and phase-specific optimizations.