Which metrics should you monitor in an LLM serving system?

Monitor client-visible TTFT, TPOT, total latency, errors, and goodput alongside server queues, token lengths, KV-cache use, and preemptions. GPU utilization alone cannot distinguish useful work from overload. Evaluate answer quality separately; a fast service can still return incorrect answers.

Start with the user-visible result

TTFT measures time from request submission to the first output token. TPOT measures the average interval between subsequent output tokens. For at least two output tokens:

TPOT=last token time−first token timeoutput tokens−1\text{TPOT} = \frac{\text{last token time} - \text{first token time}}{\text{output tokens} - 1}

Measure these endpoints at delivery of the first and last output token. Record request-completion latency separately; final metadata or connection cleanup can occur after the last token. Also record individual delivery stalls. A per-request TPOT average can hide a pause. SSE chunks may contain several tokens, so chunk intervals and token intervals are different measurements.

Report latency percentiles by relevant workload groups, such as prompt length and model. Keep timing boundaries consistent: server metrics exclude some network and client processing time.

Goodput counts requests per second that meet all defined serving SLOs. If the service completes 100 requests per second and 40 miss at least one required threshold, goodput is 60 requests per second. This measures serving performance, not factual answer quality. DistServe uses latency-constrained goodput when evaluating serving designs.

Add the metrics that explain failures

vLLM’s metrics documentation describes running and waiting requests, latency histograms, token counts, KV use, and preemptions. Bind dashboards to the deployed version.

MeasurementWhat to investigate when it rises
Waiting requests and queue timeArrival rate versus admitted processing capacity
KV-cache use and preemptionsLong sequences, active batches, and memory allocation
TTFT with stable TPOTQueueing, tokenization, prefill, or network delay
TPOT and delivery stallsDecode scheduling, memory traffic, or client buffering
Tokens and retry counts per completed requestLonger outputs or repeated work

The vLLM gauge vllm:kv_cache_usage_perc uses a fraction from 0 to 1 despite its name. Derive alert thresholds from load tests and the required recovery time. Keep a small set of quality evaluations and failure traces linked to the serving metrics.

For the broader design, read monitoring LLM systems in the guide.