How CUDA kernel fusion improves LLM inference speed
CUDA kernel fusion combines operations that would otherwise run as separate GPU kernels. It can reduce launch overhead and avoid writing an intermediate tensor to GPU memory only to read it back immediately. It helps when those costs are a meaningful part of inference latency.
Follow the intermediate data
Consider a residual addition followed by normalization. Separate kernels can write the addition result to main GPU memory and read it again for normalization. A compatible fused implementation can reuse intermediate values inside the kernel. That saves transfers; it does not remove the mathematical operations or their correctness requirements.
Fusion also applies to larger operations. DeepFusionKernel studies fusing a complete SwiGLU feed-forward computation, including its matrix multiplications. This is more than combining element-wise activation functions.
FlashNorm uses a different transformation: it incorporates RMSNorm’s learned scale into a following linear operation and defers rescaling.
Check whether fusion helped
| Compare | What an improvement would show |
|---|---|
| Kernel launches and launch gaps | Less overhead between small operations |
| HBM reads and writes | Fewer intermediate transfers |
| Register and shared-memory use | Whether the combined kernel needs more local storage |
| Spills and active blocks | Whether resource pressure reduced parallel execution |
| Full request latency | Whether the kernel change mattered to users |
The Hopper tuning guide explains the resource limits that constrain concurrently active GPU work. Combining operations can increase register use, cause spills, or limit the blocks that execute at once. A larger fused kernel can therefore be slower for some shapes.
Keep model, precision, GPU, input length, and batch size fixed. Check numerical outputs and task quality as well as time. Also compare small-batch decode with larger prefill: reducing launch overhead may matter much more for short operations than for long matrix computations.
Engineering Guide: CUDA kernels and fusion relates these methods to quantized matrix kernels.