What FlashAttention changes in transformer inference

FlashAttention computes exact attention without storing the full attention-score and probability matrices in main GPU memory. It processes smaller tiles and combines their results with an online softmax. This reduces memory traffic and intermediate storage while retaining dense attention’s quadratic arithmetic in sequence length.

Lower memory use does not make dense attention computation grow linearly.

Separate storage from arithmetic

Standard materialized attention forms scores for every visible query/key pair, normalizes those scores, and multiplies them by values. An unmasked square attention matrix contains N² entries for N tokens.

The original FlashAttention paper reorganizes that computation into tiles that fit on chip. It keeps running normalization statistics and output accumulators instead of writing all probabilities to HBM. The required attention intermediates grow linearly with sequence length for fixed head dimensions. It still evaluates the dense query/key interactions required by the attention mask.

This is exact attention as an algorithm; different floating-point execution orders can cause small numerical differences. It is not an approximate sparse-attention method that discards selected token pairs.

Implementation and workload still matter

QuestionWhy to check it
Does the backend support this GPU and precision?Kernel implementations target particular hardware
Are the head dimension, mask, and cache layout supported?Unsupported shapes can select another backend
Is attention a large part of request time?A large kernel gain can have a small overall effect
Is the workload prefill or single-token decode?Their parallel work differs

FlashAttention-2 additionally parallelizes query-sequence blocks alongside batch and heads, and reduces other execution overhead. FlashAttention-3 targets Hopper-specific execution features. Version names describe algorithms and implementations; they do not establish the backend your server actually selected.

Verify the selected backend in the pinned serving runtime. Measure prompt latency, peak memory, and end-to-end throughput at representative lengths and batch sizes. Compare identical masks and precision. A benchmark reporting kernel TFLOPS alone does not establish improvement in user-visible token latency.

Engineering Guide: FlashAttention covers successive versions and their scoped measurements.