Which GPU hardware specifications matter for LLM serving?

For LLM serving, check GPU memory capacity, memory bandwidth, compute throughput at the intended precision, and GPU interconnects. Capacity determines whether weights and request state fit. Bandwidth often limits small-batch decoding. Compute and communication matter more as prompts, batches, or distributed execution grow.

A TFLOPS figure alone cannot determine how many requests a server can handle at an acceptable latency.

Read specifications with their conditions

SpecificationWhat to ask
Memory capacityDo weights, KV cache, temporary buffers, and concurrent requests fit?
Memory bandwidthHow many bytes must each decode step read?
Tensor Core throughputWhich precision, sparsity mode, and GPU variant does the number use?
GPU-to-GPU bandwidthIs it aggregate bidirectional bandwidth or the bandwidth of one connection?
Network bandwidthWhat adapter, port count, topology, and transfer path are available?

HBM is stacked main memory used by data-center accelerators. GDDR is used by GPUs including the RTX 4090. NVIDIA lists 3.35 TB/s for H100 SXM, 4.8 TB/s for H200, and up to 8 TB/s for B200. These values describe advertised memory bandwidth, rather than measured token throughput. Sources: H100, H200, and HGX component specifications.

Separate computation from communication

An NVIDIA streaming multiprocessor, or SM, contains general compute units, Tensor Cores, and local memory resources. Tensor Cores accelerate supported matrix operations; CUDA cores handle other arithmetic. H100 SXM has 132 SMs, while the PCIe variant has 114. Even the product name needs a variant. NVIDIA’s Hopper architecture table gives the distinction.

NVLink connects supported GPUs within a system. InfiniBand connects machines. GPUDirect RDMA allows a compatible network device to transfer data directly to GPU memory without staging it in host memory. The CPU can still manage the operation. These features matter when a model is split across GPUs or its KV state moves between servers.

Record the exact GPU variant, memory, precision, server layout, and interconnect before comparing a benchmark. Then test the prompt lengths, output lengths, and concurrency you plan to serve.

Engineering Guide: GPU hardware glossary defines the additional hardware terms used in serving documentation.