Why is LLM decode memory-bound and prefill compute-bound?

LLM decode is often memory-bound because a small batch reuses each model weight only a few times before loading the next weights. Long or sufficiently batched prefill reuses weights across many prompt tokens, so matrix computation can become the limiting resource. These are workload conditions, rather than permanent properties of either phase.

Arithmetic intensity helps decide whether to reduce data movement or improve matrix execution.

Compare calculations with bytes moved

Arithmetic intensity is the number of floating-point operations performed per byte transferred from memory. The roofline model compares it with peak compute divided by memory bandwidth.

NVIDIA’s H100 SXM specifications give approximately 989 TFLOPS for dense FP16/BF16 Tensor Core computation and 3.35 TB/s memory bandwidth. The listed 1,979-TFLOPS figure assumes structured sparsity. Their ratio is about 295 FLOPs per byte. This is a theoretical bound using matching hardware specifications.

A simplified dense matrix-vector multiply performs about two operations per two-byte weight: roughly one FLOP per byte. A batch-one decoder in this regime cannot use the GPU’s peak matrix throughput. Larger batches let a loaded weight contribute to several token computations. Prefill similarly reuses weights across prompt tokens. Splitwise describes these different phase characteristics.

Choose a change from the measured limit

ObservationChange to investigateWhat can prevent a gain
Small-batch decode spends time moving weightsSupported weight quantization; larger decode batchesQuantization kernels, quality loss, latency limits
Long-context decode spends time reading attention historyFewer KV heads; supported KV quantization; efficient attentionModel architecture and cache precision support
Large prefill uses matrix units heavilyEfficient matrix kernels; supported lower-precision computeShort prompts, small batches, other operations
Many short kernels leave gaps in GPU executionFusion or reduced launch overheadAdded register/shared-memory use

Measure both first-token latency and later token intervals after each change. A larger batch can improve total tokens per second while making each request slower. The hardware ratio identifies a possible limit; a trace establishes which resource your workload actually waits for.

Engineering Guide: memory-bound versus compute-bound inference includes the roofline diagram and phase-specific optimizations.