How GPU memory hierarchy affects LLM inference speed
GPU memory hierarchy affects LLM inference because moving an intermediate tensor to main GPU memory costs more than reusing it inside a running kernel. HBM holds the large model and attention cache. Registers, shared memory, and L2 provide smaller places to reuse data during computation.
Engineers profiling inference need to distinguish memory capacity from memory traffic. A model fitting in HBM does not establish that its kernels move data efficiently.
What each memory holds
| Memory | Scope | Role in inference |
|---|---|---|
| Registers | Values used by a thread | Accumulators and intermediate calculations |
| Shared memory | Threads in one compute block | Cooperative reuse of matrix tiles and reductions |
| L2 cache | Shared across the GPU | Reuse across blocks without another main-memory fetch |
| HBM | Main GPU memory | Weights, KV cache, inputs, outputs, and larger intermediates |
For H100 SXM, NVIDIA documents 80 GB HBM3 at 3.35 TB/s, 50 MB L2, and up to 228 KB shared memory per streaming multiprocessor. The amount available to one thread block is smaller and depends on allocation settings. These are specifications for this GPU variant, rather than limits shared by every H100. See the H100 specification and Hopper tuning guide.
Why an intermediate tensor matters
A sequence of separate kernels may write an intermediate result to HBM and read it back in the next kernel. Fusion can keep that result in registers or shared memory. FlashAttention applies related data-reuse principles to attention, avoiding large score matrices in HBM.
On-chip storage is limited. A kernel that uses too many registers or too much shared memory can reduce the number of blocks that run concurrently. Register spills can add memory traffic. Therefore, more fusion does not guarantee faster execution.
When investigating a change, compare HBM bytes transferred, register use, shared-memory use, spills, and the number of active blocks. Keep the model, tensor shapes, precision, and GPU variant fixed. Then measure the full request: a faster attention kernel may account for only a small share of total latency.
Engineering Guide: GPU memory hierarchy shows where these memories are located on an H100 SXM.