Continuous vs static batching: how LLM scheduling works

Static batching groups requests for an inference run. Continuous batching updates the active request set between generation iterations: completed requests leave and eligible waiting requests enter. This lets a server reuse capacity without waiting for every original request to finish.

For serving engineers, the benefit depends on variable response lengths, scheduler policy, memory capacity, and latency requirements.

Compare unequal output lengths

Suppose three requests need 10, 40, and 100 output tokens. In a simple fixed-batch generation loop that does not replace finished requests, capacity occupied by the first two becomes available before the third finishes. The server still waits for the batch’s longest response before starting another complete batch.

An iteration-level scheduler can remove each finished request and admit another. Orca introduced an LLM serving design using this scheduling approach. Admission still needs free KV-cache blocks and a compatible iteration budget. Continuous batching does not mean every queued request starts immediately.

New prompts also need prefill. Chunked prefill can share iteration budgets with running decodes. The scheduler’s admission and priority rules therefore affect both time to first token and later token intervals.

Compare the scheduler and the full runtime

Anyscale’s 2023 experiment reported about 4x for FasterTransformer, 8x for its continuous-batching configurations, and 23x for vLLM against its naive baseline. It used OPT-13B, an A100-40GB, 1,000 requests, 512-token inputs, and an exponential output distribution with a mean of 128 tokens. These are results for entire implementations in that setup, rather than the isolated effect of one scheduling change.

For your comparison, record:

CheckWhy it matters
Request arrival and output-length distributionsDetermine unused capacity and queue pressure
Maximum active sequences and token budgetBound work in each iteration
KV allocation and preemptionCan stop new admissions
TTFT and per-request generation latencyExpose the cost to individual requests
Success rate and goodputShow whether extra work meets the service target

Inspect short and long requests separately. Higher total throughput can coexist with worse latency for one request group. Choose scheduling settings from the measured service requirements.

Engineering Guide: continuous batching explains the allocation and scheduling mechanisms together.