What is Flash-Decoding, and when does it speed up LLMs?
Flash-Decoding parallelizes decode attention across the cached key/value sequence. Different GPU blocks process different parts of a long history, then combine partial attention outputs with a numerically correct normalization. It can help small-batch generation when a long cache provides more work than the existing attention schedule executes efficiently.
It accelerates an attention operation. The resulting serving speedup also depends on time spent elsewhere in the model.
Why decode needs a different schedule
During prefill, many query positions can run in parallel. During ordinary decode, each active sequence contributes one new query position. A small batch can therefore provide too few independent blocks to use the GPU effectively, even though each query must read a long history.
Flash-Decoding creates additional work units by splitting that history. Each unit computes a partial output and a log-sum-exp normalization value. A reduction combines them into the output that attention over the complete history requires. See the Stanford authors’ Flash-Decoding explanation.
For a short history, the extra partial-result and reduction work may outweigh the benefit. A larger batch can already supply enough parallel work. Neither case guarantees a gain from further sequence splitting.
Interpret the reported result
The authors reported up to eightfold end-to-end speedup for their CodeLlama-34B batch-one evaluation on four A100 GPUs, covering sequence lengths from 512 to 64K. That comparison used their selected baseline and execution setup. It does not predict an eightfold improvement over a current server with an already optimized decode backend.
When testing your server:
- verify that its selected attention backend actually uses the method;
- compare the same model, GPU layout, cache precision, and sequence lengths;
- test both batch one and realistic concurrency;
- record attention-kernel time separately from the complete decode step;
- measure output-token latency and numerical correctness.
A faster attention kernel may leave weight reads, feed-forward computation, or communication as the largest remaining cost. Prefer the end-to-end result when deciding whether to change the serving configuration.
Engineering Guide: Flash-Decoding connects sequence splitting to the prefill/decode distinction.