What causes LLM serving failures, and how do you fix them?

LLM serving failures often come from insufficient memory, more concurrent work than the service can handle, or retries that repeat expensive requests. Identify which resource or stage fails before changing the runtime. A model loading successfully does not prove that it can handle the intended traffic.

Match the symptom to evidence

SymptomFirst evidence to inspectChange to test
GPU out-of-memory errorWeights, KV allocation, activations, request lengthsSmaller active batch, shorter limits, or supported quantization
Latency rises without errorsKV use, preemptions, waiting requestsLower admission concurrency or more measured capacity
Slow first tokenQueue, tokenization, prefill, network timingsFix the slow stage; test chunked prefill if phases interfere
Load increases after timeoutsRetry counts and work still runningBounded retries and cancellation of abandoned work

Weight storage and KV storage are separate. A 70B FP16 model needs roughly 140 GB of raw weight values. Under a conventional full-attention cache layout, one Llama 3.1 70B sequence at 128K tokens adds 40 GiB of FP16 KV data before runtime overhead. Neither number alone is a complete memory requirement.

Check preemption before increasing the queue

When KV space is insufficient, a serving engine may preempt requests. vLLM V1 normally recomputes preempted work, which adds latency even if the API eventually succeeds. Correlate preemption counts with prompt lengths, active sequences, and KV use.

Reduce admitted work or change the tested memory allocation. Increasing GPU memory utilization requires confirming that other allocations still fit. Moving reusable KV data to CPU or disk also adds transfers; it does not guarantee that an oversized active request fits.

Keep retries bounded

Retry transient failures within a deadline and retry budget. Use backoff with jitter, and assign retries to one layer so gateway and client retries do not multiply. Google’s SRE guidance explains how retries can worsen overload.

Propagate cancellation when the caller leaves, verify that generation stops, and load-test this behavior. More queue capacity delays rejection but does not create processing capacity.

Read the guide’s LLM serving failure modes for the surrounding memory and scheduling concepts.