How do you plan GPU capacity and autoscaling for LLMs?
Size the fleet from measured throughput at the required latency, then verify that each replica can hold the intended active sequences. Use real prompt and output-length distributions. Round up complete serving replicas, including every GPU required by tensor or pipeline parallelism.
Check memory before throughput
Under a conventional full-attention cache layout, Llama 3.1 70B needs 1.25 GiB of FP16 KV data for a 4,096-token sequence. A hypothetical 40 GiB KV budget therefore holds at most 32 such sequences before allocator and runtime overhead. At 131,072 tokens, one sequence uses the full 40 GiB.
Count input and generated tokens together. The formula changes for architectures such as MLA and sliding-window attention, and per-device allocation depends on sharding. This is a memory bound, not a prediction of the batch size that meets latency targets. The guide’s KV-cache calculation shows the assumptions and runnable arithmetic.
Calculate whole replicas
Suppose a two-GPU replica has been measured to sustain 1,200 output tokens per second at the required TTFT and TPOT on a fixed workload. Expected peak demand is 4,000 output tokens per second. A hypothetical 1.3 safety multiplier gives:
Five two-GPU replicas require ten GPUs. These figures illustrate the calculation; they are not hardware benchmark results. Re-measure when the model, context distribution, topology, or SLO changes. A multiplier of 1.3 means 30% more throughput than demand before rounding, rather than 30% unused capacity.
Scale before deadlines fail
Combine queue depth and waiting time with KV use, preemptions, and SLO-compliant goodput. GPU utilization can remain high after the service becomes overloaded. vLLM documents these serving metrics; derive thresholds from load tests.
Measure image pulls, weight loading, model initialization, and warmup before choosing autoscaling thresholds. KEDA’s deployment scaling controls can scale workloads toward zero, but they do not remove model startup latency. Keep ready capacity when requests cannot wait for startup.
Read capacity planning and autoscaling in the guide for the surrounding hardware and admission decisions.