Continuous vs static batching: how LLM scheduling works
Static batching groups requests for an inference run. Continuous batching updates the active request set between generation iterations: completed requests leave and eligible waiting requests enter. This lets a server reuse capacity without waiting for every original request to finish.
For serving engineers, the benefit depends on variable response lengths, scheduler policy, memory capacity, and latency requirements.
Compare unequal output lengths
Suppose three requests need 10, 40, and 100 output tokens. In a simple fixed-batch generation loop that does not replace finished requests, capacity occupied by the first two becomes available before the third finishes. The server still waits for the batch’s longest response before starting another complete batch.
An iteration-level scheduler can remove each finished request and admit another. Orca introduced an LLM serving design using this scheduling approach. Admission still needs free KV-cache blocks and a compatible iteration budget. Continuous batching does not mean every queued request starts immediately.
New prompts also need prefill. Chunked prefill can share iteration budgets with running decodes. The scheduler’s admission and priority rules therefore affect both time to first token and later token intervals.
Compare the scheduler and the full runtime
Anyscale’s 2023 experiment reported about 4x for FasterTransformer, 8x for its continuous-batching configurations, and 23x for vLLM against its naive baseline. It used OPT-13B, an A100-40GB, 1,000 requests, 512-token inputs, and an exponential output distribution with a mean of 128 tokens. These are results for entire implementations in that setup, rather than the isolated effect of one scheduling change.
For your comparison, record:
| Check | Why it matters |
|---|---|
| Request arrival and output-length distributions | Determine unused capacity and queue pressure |
| Maximum active sequences and token budget | Bound work in each iteration |
| KV allocation and preemption | Can stop new admissions |
| TTFT and per-request generation latency | Expose the cost to individual requests |
| Success rate and goodput | Show whether extra work meets the service target |
Inspect short and long requests separately. Higher total throughput can coexist with worse latency for one request group. Choose scheduling settings from the measured service requirements.
Engineering Guide: continuous batching explains the allocation and scheduling mechanisms together.