MHA vs MQA vs GQA: how attention changes KV memory use
MHA gives each query head its own key and value head. MQA shares one key/value head across all query heads. GQA shares a key/value head within each group of query heads. Fewer KV heads reduce the conventional attention cache and the data that decode reads from it.
When estimating cache size, check num_key_value_heads. Query-head count alone can substantially overestimate its cache.
Compare head sharing directly
Assume eight query heads and identical head dimension, layer count, sequence length, and cache precision:
| Attention | Query heads | KV heads | Cache relative to MHA |
|---|---|---|---|
| Multi-head attention, MHA | 8 | 8 | 100% |
| Multi-query attention, MQA | 8 | 1 | 12.5% |
| Grouped-query attention, GQA | 8 | 2 | 25% |
A KV head means one key head and one value head. Query projections remain separate. Shazeer’s MQA paper and Ainslie et al.’s GQA paper define these changes.
For Llama 3 70B, 64 query heads share eight KV heads: an eightfold reduction in raw KV payload against the otherwise identical model with 64 KV heads. Llama 3.1 405B has 128 query heads and eight KV heads, giving a sixteenfold reduction by the same comparison. Meta’s architecture table supplies the counts.
What the reduction does not promise
An eightfold smaller KV cache does not mean eightfold less total GPU memory or eightfold faster generation. Weights, activations, kernel execution, and communication remain. Distributed implementations may also replicate KV heads across partitions.
The GQA paper found a useful quality/speed trade-off in its tested T5-family models. Changing a trained MHA checkpoint to GQA is an architectural change that needs adaptation and evaluation; it is not an allocation setting.
Check the checkpoint’s attention configuration, use KV-head count in the cache formula, and benchmark the server at the context lengths and concurrency you need. Keep task quality in that comparison when choosing between checkpoints.
Engineering Guide: GQA and MQA illustrates the head-sharing patterns.