When does speculative decoding make LLM serving faster?
Speculative decoding can accelerate generation when proposing several candidate tokens is cheap and the target model accepts enough of them. The target verifies the candidates together, potentially returning several output tokens per round. Extra draft work can instead make serving slower when acceptance is low or the target already runs efficiently in large batches.
Benchmark the complete generation loop, including drafting and rejection, before choosing it for a service.
Understand a verification round
A draft model proposes K tokens. The target scores the candidate positions in a batched forward pass. Modified rejection sampling accepts candidates from left to right using both distributions. After the first rejection, it samples a correction from the residual target distribution and discards later candidates. If every candidate passes, it can also sample one additional target token.
Chen et al.’s speculative-sampling algorithm preserves the target distribution under the algorithm’s assumptions, subject to hardware numerics. Merely comparing candidate tokens with target-selected tokens is not the same sampling algorithm. Check the selected implementation’s sampling behavior and supported settings.
The paper reported 2–2.5x decoding speedup for its tested Chinchilla-70B target with a 4B draft, batch one, and K=4 on TPU v4. A historical result is evidence that the mechanism can help, rather than a forecast for another model pair.
Check what determines the gain
| Measure | Interpretation |
|---|---|
| Accepted tokens per verification round | How much output amortizes target execution |
| Draft time | Cost of proposing candidates |
| Verification and correction time | Target work and rejected-candidate overhead |
| Throughput and latency by concurrency | Whether the gain survives realistic batches |
| Task quality and sampling settings | Whether the selected method preserves intended behavior |
Draft-model methods are one option. EAGLE-3 combines several levels of target features and predicts draft tokens directly; other approaches use added prediction heads or prompt matching. Their acceptance rates and execution costs differ.
Test representative code, prose, structured outputs, and context lengths. Start with a small draft length, then vary it with concurrency. More candidates do not help if their proposal and verification cost exceeds the work avoided.
Engineering Guide: speculative decoding explains the acceptance sequence and method variants.