What causes LLM serving failures, and how do you fix them?
LLM serving failures often come from insufficient memory, more concurrent work than the service can handle, or retries that repeat expensive requests. Identify which resource or stage fails before changing the runtime. A model loading successfully does not prove that it can handle the intended traffic.
Match the symptom to evidence
| Symptom | First evidence to inspect | Change to test |
|---|---|---|
| GPU out-of-memory error | Weights, KV allocation, activations, request lengths | Smaller active batch, shorter limits, or supported quantization |
| Latency rises without errors | KV use, preemptions, waiting requests | Lower admission concurrency or more measured capacity |
| Slow first token | Queue, tokenization, prefill, network timings | Fix the slow stage; test chunked prefill if phases interfere |
| Load increases after timeouts | Retry counts and work still running | Bounded retries and cancellation of abandoned work |
Weight storage and KV storage are separate. A 70B FP16 model needs roughly 140 GB of raw weight values. Under a conventional full-attention cache layout, one Llama 3.1 70B sequence at 128K tokens adds 40 GiB of FP16 KV data before runtime overhead. Neither number alone is a complete memory requirement.
Check preemption before increasing the queue
When KV space is insufficient, a serving engine may preempt requests. vLLM V1 normally recomputes preempted work, which adds latency even if the API eventually succeeds. Correlate preemption counts with prompt lengths, active sequences, and KV use.
Reduce admitted work or change the tested memory allocation. Increasing GPU memory utilization requires confirming that other allocations still fit. Moving reusable KV data to CPU or disk also adds transfers; it does not guarantee that an oversized active request fits.
Keep retries bounded
Retry transient failures within a deadline and retry budget. Use backoff with jitter, and assign retries to one layer so gateway and client retries do not multiply. Google’s SRE guidance explains how retries can worsen overload.
Propagate cancellation when the caller leaves, verify that generation stops, and load-test this behavior. More queue capacity delays rejection but does not create processing capacity.
Read the guide’s LLM serving failure modes for the surrounding memory and scheduling concepts.