LLM throughput vs goodput: which metric should you use?

LLM throughput measures how much work a server completes per second. Goodput measures the rate of work that also meets defined service requirements. For serving decisions, use output-token throughput together with successful requests per second that meet your first-token and generation-latency targets.

A server that produces more tokens while requests time out may have higher throughput and lower useful capacity.

Define a qualifying request

Before testing, define the conditions a request must meet. For example: successful completion, TTFT below 500 ms, and average TPOT below 50 ms. Then count qualifying completions during the measured interval:

request goodput = qualifying completed requests / measured seconds
output throughput = generated output tokens / measured seconds

These are an illustrative SLO and two explicit metric definitions. If 100 requests complete in ten seconds but only 60 meet the requirements, completion throughput is ten requests per second and goodput is six. Also report output-token throughput, since response lengths can differ.

DistServe uses goodput to evaluate serving under TTFT and token-generation requirements. Always state whether your unit is qualifying requests or qualifying tokens: the word alone does not define the denominator or eligibility rules.

Keep the workload comparable

RecordWhy it affects the result
Input and output lengthsPrefill work, cache size, and decode duration differ
Model, precision, GPU, and runtime revisionThey change the execution cost
Offered requests per secondDetermines queue pressure
Concurrent requestsChanges batching and memory use
Cache policyRepeated prefixes can reduce prefill work
Errors and cancellationsSuccessful completion is part of useful capacity

At small decode batches, adding requests can improve weight reuse. Larger batches or increasing offered load can eventually raise latency or exceed memory, computation, or communication limits. The curve depends on the workload; it is not guaranteed to scale linearly.

Increase offered load in steps and plot throughput, goodput, P99 latency, errors, and queue length. Choose a load that meets the requirements with room for traffic variation. The Anyscale continuous-batching experiment shows why a comparison needs a stated request distribution.

Engineering Guide: throughput and the latency tradeoff explains the serving trade-off in context.