GPU Selection for LLMs: How Much GPU Memory Do I Need?
Estimate weights, KV cache, and runtime memory before comparing GPU speed. Then benchmark bandwidth, compute, and interconnect with the intended request distribution. A model whose weights fit may still exceed memory when several long requests run together.
Use the required concurrency and sequence lengths in the memory calculation.
Build the memory estimate
A 7B-parameter model stored in FP16 has about 14 GB of raw weight values. INT4 reduces those values to about 3.5 GB, but scales, unquantized tensors, and runtime allocations add memory. The artifact’s actual allocation is the relevant number.
Next estimate KV-cache memory for the architecture, cache precision, active sequence lengths, and batch size. Finally include temporary activations, buffers, allocator overhead, and a measured reserve.
For a simple planning example, 14 GB of weights plus a measured 4 GB KV allocation already require 18 GB before other allocations. That leaves little room on a 24 GB card if runtime overhead or longer requests consume the rest. It does not establish that every 7B model has a 4 GB cache.
Bandwidth and GPU links change the comparison
| GPU variant | Published memory | Published memory bandwidth |
|---|---|---|
| H100 SXM | 80 GB HBM3 | 3.35 TB/s |
| H200 SXM | 141 GB HBM3e | 4.8 TB/s |
| B200 SXM | 180 GB HBM3e | Up to 8 TB/s |
These values come from NVIDIA’s HGX reference architecture. They describe particular variants, not interchangeable specifications for every card with the same family name.
At low-batch decode, reading weights and attention state can make memory bandwidth more relevant than peak matrix throughput. Larger batches, long prefills, and optimized kernels can change the limiting resource. When a model needs multiple GPUs, include collective-communication time and the available links between devices.
Compare actual cost per successful request at the same quality and latency target. Include required GPU groups and unused capacity; current rental price alone does not show serving efficiency.
The GPU-selection section of the LLM Engineering Guide compares additional devices. Use KV-cache memory planning to estimate the variable allocation.