llamaperf
← All articles

Reading benchmarks

Speculative decoding: when a draft model speeds you up

By llamaperf · · 5 min read

Quick answer

Speculative decoding lets a small, fast model guess the next few tokens so the big model can check them in one pass. When the guesses are right, text comes out faster and the output matches what the big model would have written. It pays off most for one user and predictable text. If you share a speed figure, say whether it was on, because results with and without it can't be compared.

How speculative decoding works

A big model spends most of each step reading its weights from memory. Checking several tokens in one step costs about the same as writing one, so a cheap guess can save time.

A small model guesses, the big one checks

The draft model proposes a few tokens. The main model scores them all at once, keeps the ones it agrees with and writes its own token where the draft went wrong. When the draft guesses well, you get several tokens for roughly the price of one. When it guesses badly, you've paid a little extra for nothing.

You don't trade away quality for this. The vLLM documentation puts it plainly:

Speculative decoding sampling is theoretically lossless up to the precision limits of hardware numerics.

From the vLLM speculative decoding docs

MTP does the same thing from inside the model

Some models ship with extra prediction heads trained to guess the next few tokens themselves, which is usually called multi-token prediction or MTP. It works on the same principle without a second model file. For reporting purposes it counts as speculative decoding too.

When a draft model pays off

The gain depends on how often the draft is right and on how busy the machine is.

Predictable text gains the most

Code and structured output give the draft lots of easy wins, and so does text that repeats your input. Creative writing at a high temperature gives it fewer. That's why two people with the same card and the same pair of models can report very different speedups. Their prompts differ.

A busy server gains less

The vLLM docs describe speculative decoding as a way to

reduce inter-token latency under medium-to-low QPS (queries per second), memory-bound workloads.

From the vLLM speculative decoding docs

A home setup with one user is exactly that case. When many requests share the GPU, the spare compute that made guessing cheap is already in use, so the benefit shrinks. So if you share your model with others, test speculation again under that load before you keep it on.

Turning it on in llama.cpp

llama-server takes a draft model with -md (also spelled --spec-draft-model), as listed in the server README. The draft has to use the same tokenizer as the main model, which in practice means a small model from the same family.

llama-server -m main-model.gguf -md small-draft.gguf --spec-draft-n-max 3

How many tokens to draft

The same README sets --spec-draft-n-max to 3 by default. Longer drafts win more when the guesses are good and waste more when they aren't. Try a couple of values on your own prompts and keep the one that's fastest end to end. Flag names in llama.cpp change between releases, so check llama-server --help on your build before you copy a command.

Check the gain yourself

Run the same prompt with the draft on and off, with the same output limit, and compare the time to a finished answer. If the gain is small, try a smaller draft model, since a draft that's too slow eats its own savings. Our calculator estimates plain decoding speed only, so a speculative run that beats it isn't a sign the estimate is broken.

Say whether it was on when you share a number

A tokens-per-second figure with speculation on measures a different thing from one without it. Mixing them makes a card look faster or slower than it is.

Why the two numbers can't share a median

Say one person reports 40 tokens per second with plain decoding and another reports 70 with a draft model, on the same card and model (made-up figures). Averaged together, they describe a setup nobody ran. Anyone comparing cards on the leaderboard or in a head-to-head comparison needs to know which kind of number they're reading.

What to include in your report

Write whether speculation was on and which method you used. If it was a draft model, add its exact file and the draft length. Add a plain-decoding run on the same prompt if you can, because the pair is far more useful than either number alone. Our benchmark-reading guide lists the other fields worth recording, and the submission form has room for all of them.

Frequently asked questions

How does speculative decoding work?

A small draft model guesses a few tokens ahead and the main model checks them all in one pass. Correct guesses are kept, so you get several tokens for roughly the cost of one step of the big model.

Does speculative decoding lower output quality?

No. The main model still decides every token, and the vLLM documentation describes the sampling as theoretically lossless up to the precision limits of the hardware.

Which draft model should I use?

Use a small model from the same family so it shares the tokenizer. Try the smallest one first, since a draft that runs slowly cancels out the time it saves.

Is speculative decoding worth it when offloading to CPU?

Test it on your setup. When part of the model sits in system RAM, each step of the main model is slower, so each good guess saves more time. A slow draft can also eat the gain, so compare runs with it on and off.

What is the difference between MTP and a draft model?

Both guess tokens ahead for the main model to check. A draft model is a separate smaller file. MTP uses extra prediction heads trained into the main model itself, so there's no second model to load.

Why do speculative decoding speedups vary so much between reports?

The gain depends on how often the draft guesses right, and that depends on the prompt, the sampling settings and the draft model. Predictable text such as code gains more than creative writing.