Chapter 5 · Inference engines
Speculative decoding
5.4

Speculative decoding

Autoregressive decoding is serial by construction: generating K tokens takes K sequential runs of the model, each one waiting on the last token before it can start. But because a decode step is memory bound rather than compute bound, as the previous chapter established, the GPU is usually sitting on spare arithmetic capacity while it waits for weights to stream in. Speculative decoding spends that spare capacity on parallel verification instead of serial generation. A small, fast proposes several tokens ahead, autoregressively, and the large then scores all of those candidates in a single parallel forward pass, the same cost as generating just one token normally.

The verification step is not a simple accept-or-reject on whether the two models agree. It uses a modified rejection sampling scheme: a draft token is kept whenever the target model would have assigned it at least as much probability as the draft model did, and when the target assigns it less probability, the token is only kept with probability equal to that ratio; a rejected token is replaced by resampling from a distribution built out of the leftover probability mass. That correction is what makes the method exact rather than an approximation: the sequence that comes out has exactly the distribution the target model would have produced sampling on its own, regardless of what draft model was used or how often it was wrong. What varies with the draft model’s quality is the , since every rejection means one parallel scoring pass produced only one or a few tokens instead of many.

Both papers that introduced this measured real speedups without retraining or changing the target model at all. Leviathan et al. report a 2 to 3 times latency improvement sampling from an 11 billion parameter T5-XXL model, walltime-tested against the standard T5X implementation. In a separate worked example, a 38-token sentence was produced from only 9 serial runs of a 97 million parameter target model paired with a 6 million parameter approximation model. Chen et al. report a 2 to 2.5 times decoding speedup on the 70 billion parameter Chinchilla model in a distributed serving setup, using the same draft-then-verify structure under the name speculative sampling.

One round of draft-then-verify. The draft model proposes five tokens one at a time; the target model scores all of them in a single forward pass. The first three are accepted, the fourth is rejected and replaced by a token resampled from the target's leftover probability mass, and everything after the rejection is discarded. One target-model step produced four tokens instead of one.Illustrative numbers

The draft model itself is the operational weak point of this design: a separate model has to be acquired and maintained, and the Medusa paper names that as the obstacle impeding adoption. Medusa removes it. Instead of a second model, it adds extra on top of the backbone’s last hidden states, each a single feed-forward layer with a residual connection predicting a token several positions ahead. The heads emit multiple top predictions per position, which are assembled into several candidate continuations and verified simultaneously in one decoding step through a tree-based attention mechanism, a simple adjustment to the attention mask. Fine-tuning only the heads on a frozen backbone, called Medusa-1, keeps the acceleration lossless and reaches over 2.2 times speedup; training the heads and backbone together, Medusa-2, raises that to 2.3 to 2.8 times.

EAGLE keeps a small drafting network but moves the drafting to a different level of the model. Its starting observations are that autoregression at the feature level, the second-to-top layer of the transformer, is more straightforward than at the token level, and that the inherent uncertainty in feature-level autoregression is what constrains its performance. EAGLE resolves that uncertainty by feeding the draft network a token sequence advanced by one time step alongside the features, which makes precise feature prediction possible with minimal overhead. On LLaMA2-Chat 70B this reaches latency speedups of 2.7 to 3.5 times and doubles throughput, while still maintaining the distribution of the generated text, the same exactness guarantee the original draft-and-verify schemes provide.

Turning it on, and knowing whether it paid#

In an engine the draft-and-verify pair collapses into a nested dictionary. vLLM’s offline configuration for a draft model speculating five tokens at a time is the same LLM constructor with one more argument:

draft model speculation, from the vLLM speculative decoding documentation
from vllm import LLM, SamplingParams

prompts = ["The future of AI is"]
sampling_params = SamplingParams(temperature=0.8, top_p=0.95)

llm = LLM(
    model="Qwen/Qwen3-8B",
    tensor_parallel_size=1,
    speculative_config={
        "model": "Qwen/Qwen3-0.6B",
        "num_speculative_tokens": 5,
        "method": "draft_model",
    },
)
outputs = llm.generate(prompts, sampling_params)

The target is the outer model, the draft an inner one from the same family and roughly a thirteenth of its size, and num_speculative_tokens is how far ahead it proposes each round. The call site does not move: generate takes the same prompts and the same SamplingParams, because the guarantee of speculative decoding is that the output distribution is unchanged. Online, the identical dictionary goes to vllm serve as --speculative-config, and clients see nothing. Not every method needs a second checkpoint: the n-gram method proposes tokens by matching n-grams already in the prompt, so its configuration is a method name, a token count, and a lookup window rather than a model path.

The choice among methods is a load question, and vLLM’s documentation frames it that way rather than by peak speedup. It describes speculative decoding as a way to reduce inter-token latency under medium-to-low queries per second on memory-bound workloads, exactly the regime this chapter opened with. Its documentation then separates model-based methods such as EAGLE, MTP, draft models and MLP speculators, the group it presents as best for latency, from n-gram and suffix decoding, described as modest speedups that do not increase workload during peak traffic. That last clause is what the papers’ headline multipliers do not carry: drafting is extra work, and at high load the batch is already large enough that there is less idle arithmetic to spend on it.

Acceptance rate stops being an abstraction the moment an engine reports it. Started with per-request speculative decoding metrics set to summary, a vLLM server attaches a speculative_decoding object to each response carrying mean_acceptance_length, draft_acceptance_rate, an acceptance histogram, and the raw counts behind them, alongside server-aggregated figures on the metrics endpoint. The sample response in the documentation reads as a diagnosis: 129 draft tokens across 43 speculation steps yielded 10 accepted, a draft acceptance rate of about 0.078 and a mean acceptance length of about 1.23. That configuration is paying for three drafts per step and getting roughly one token back, and the counters say so before any latency measurement has to be interpreted.

When speculation pays#

Calling that configuration bad is a judgement, and the original paper supplies the arithmetic that makes it a measurement. Leviathan et al. model a round of speculation as a capped geometric variable: if each proposed token is accepted independently with probability alpha, and the draft proposes gamma tokens per round, then the expected number of tokens one round produces is (1 - alpha^(gamma+1)) / (1 - alpha). Expanded, that quotient is the geometric sum 1 + alpha + alpha^2 and so on through alpha^gamma: one token the target model would have emitted anyway, plus one more for each draft position that survives verification. Alpha is not a fitted constant either. The paper derives it as the expected overlap between the two models’ next-token distributions, the expectation of min(p, q) taken position by position, which is why a draft can score well by agreeing with the target wherever the target is confident without resembling it anywhere else.

Acceptance alone still decides nothing, because drafting is not free. The paper names its price: c, the cost coefficient, is the ratio between the time for a single run of the draft model and the time for a single run of the target. With it the expected walltime improvement is the expected token count divided by gamma * c + 1, the extra denominator being the gamma serial draft runs each round now has to pay for. Two consequences matter. Alpha is an intrinsic property of the models and the task, while c depends on the hardware configuration and software implementation details, so the same pair of checkpoints can pay on one deployment and lose on another with no change to either model. And the condition for any speedup at all is a single comparison: if alpha exceeds c, some gamma yields an improvement, and that improvement is at least (1 + alpha) / (1 + c). Below that line no draft length rescues the configuration, which is the difference between tuning num_speculative_tokens and abandoning the draft model.

The counters connect to the formula directly, without any need to estimate alpha. vLLM defines mean_acceptance_length as the mean tokens emitted per verification step including the bonus token, which is exactly the quantity the expectation computes. So the reader’s test is to divide the reported mean acceptance length by gamma * c + 1 and ask whether the result clears one. For the documentation’s own sample that is 1.23 against three drafts a step, and in the paper’s experiments, where the draft was typically a couple of orders of magnitude smaller than the target, c was always less than 0.05 and often negligibly close to zero. Even taking 0.05 as a ceiling puts the denominator at 1.15 and leaves a quotient barely above one, so the configuration is not merely unimpressive: it is close to the line where drafting stops being worth doing at all. Choosing an approximation model around two orders of magnitude smaller than the target is the paper’s own answer to the same trade, balancing alpha against c.

The formula also explains the load advice rather than just restating it. Speculation does not reduce arithmetic; it increases it, by a factor the paper works out as (1 - alpha)(gamma * c + gamma + 1) / (1 - alpha^(gamma+1)) when c is measured in operations rather than time, with the plain summary that if alpha is low the increase in the number of arithmetic operations is high, and vice versa. What it reduces is memory traffic: the target model’s weights and KV cache are read once per round instead of once per token, so the memory accesses for reading them shrink by the same factor as the token count grows. That is the whole bargain of the technique stated in two directions, and it is why vLLM warns against the method during peak traffic. At low load the wasted arithmetic was idle anyway; at high load a large batch already wanted it.