How to measure LLM latency using TTFT, TPOT, and P99
Time to first token, or TTFT, measures how long a request waits for its first generated token. Time per output token, or TPOT, measures the average interval between later tokens. Measure both: a server can start quickly and generate slowly, or make users wait before streaming quickly.
Engineers comparing servers should use the same timing boundaries and report latency by prompt length, output length, and load.
Calculate each request separately
Record request submission, first generated token, final generated token, and output-token count. For a response containing at least two tokens:
TTFT = first_token_time - request_start_time
TPOT = (last_token_time - first_token_time) / (output_tokens - 1)
If submission occurs at 0 seconds, the first token arrives at 0.3 seconds, and the twentieth arrives at 1.25 seconds, TTFT is 300 ms and average TPOT is 50 ms. TPOT is undefined for a one-token response.
Client-visible TTFT can include networking, queueing, tokenization, and prefill. A server-side prefill timer excludes several of those costs. Label both measurements rather than comparing them as equivalent. The prefill pass commonly supplies the logits used to generate the first token; see Splitwise’s inference-phase description.
Inspect the latency distribution
P50 is the median request latency. P99 is the value below which 99% of measured requests fall. Calculate the percentiles over per-request measurements. Averaging every token interval across all requests weights long responses more heavily and answers a different question.
Average TPOT also hides pauses between particular tokens. Record individual inter-token intervals when diagnosing streaming stalls. An HTTP chunk may carry several tokens, so chunk-arrival intervals alone are not exact token-generation timings.
For a historical reference, MLPerf Inference v5.0 specified P99 TTFT of 450 ms and P99 TPOT of 40 ms for its Llama 2 Chat 70B interactive benchmark. Those limits define that benchmark. Choose your product’s targets through user testing, then report the offered load at which the server meets them.
Engineering Guide: TTFT, TPOT, and percentiles connects these measurements to queueing and inference phases.