llamaperf
← All articles

Reading benchmarks

Why the same GPU can show very different local LLM speeds

By llamaperf · · 6 min read

Quick answer

The GPU name is only part of a speed result. Before you read anything into a gap, match the model build, quantization, engine, GPU count and workload. Reading the prompt and writing the answer are timed separately, and a server total can add up many users at once.

Work out which number you're looking at

A post that says "40 tokens per second" skips a question: which tokens? Reading your prompt and writing the answer are separate stages, and the first one is usually far faster.

Prompt processing and generation are separate numbers

llama.cpp's own benchmark tool reports them apart. The example table in the llama-bench README shows a 7B Qwen2 build at Q4_K_M processing a 512-token prompt at 7,340 tokens per second and generating 128 tokens at 120.6 tokens per second, on the same machine. That's the same card and the same file with a gap of about 60 times. Compare one person's first number with another person's second and you'll think a GPU is broken.

Servers add a third kind of number

Serving engines split the stages too. vLLM's metrics reference lists prompt tokens, generation tokens and time to first token as separate metrics. A server with eight people using it can also report the combined output of all of them, which is far higher than what any one of them sees.

Before you compare two results, find out which of these each one measured. If the post doesn't say, treat the number as an observation with missing context. Our performance guide walks through the difference.

Line the reports up side by side

Copy each setup into a small table. Most apparent disagreements go away once the rows sit next to each other.

The fields worth comparing

FieldWhy to check it
Model and variantSimilar names can be different downloads
QuantizationDifferent weights, sometimes a different code path
Engine and versionThe software is part of the setup
GPU countA card name doesn't say how many were used
Input and output lengthThe workloads may differ a lot
ConcurrencyA server total isn't one user's experience
Cache and warmup stateReused work and startup both change timing

Quantization catches people out the most. In the Llama 3.1 8B table in the llama.cpp quantize README, the 3.42 GiB IQ3_S build generates 69.31 tokens per second and the bigger 4.36 GiB Q4_K_S build generates 76.71. A smaller file can be slower.

Leave unknowns blank

Leave a field empty when the source doesn't give it. Filling in the usual default makes the table look complete and hides the fact that you're guessing. The compare page puts two cards side by side from community reports, and it's a quick way to see which fields differ.

Reproduce the gap yourself

If you own the card, you can settle it. Fix the prompt and the output limit, and save your starting configuration before you touch anything.

Repeat every run

Run each setup several times and keep every result, slow ones included. llama-bench does this for you by default:

"Each test is repeated the number of times given by -r, and the results are averaged." (llama-bench README)

Its default is five repetitions, and the README says the output includes the standard deviation, so you can see how much the runs moved around.

Change one setting at a time

If you swap the model file and the engine together and it gets faster, you have a faster setup and no idea why. That's fine when you just want a daily driver. It won't tell anyone which change mattered.

Also note whether the model was already loaded and whether you sent the exact same prompt twice. Some engines reuse work from an identical prompt, so a repeat can come back much quicker than anything a real user would see.

Post a number other people can use

A speed figure is only useful to others if they can see how you got it.

What to include

Share the setup with the number. Say how you measured it, and whether it's one run or the median of several. When results bounce around, a range tells readers more than your best run.

If you're comparing two setups, say which fields match and which don't. When you can't control the test, "these are two community reports with different conditions" is a fair conclusion. Calling one GPU faster because its number is bigger is the exact mistake this checklist is here to stop.

Where to send it

Our report template keeps the important fields together. Fill it in, then submit your result. A slow result with full settings helps the next reader more than a record with none. If you're still choosing what to run, start with how to choose a local LLM for your hardware.

Frequently asked questions

Why is my smaller GGUF getting fewer tokens per second than a bigger one?

Different quant types use different code paths, and some small ones decode more slowly. In llama.cpp's own Llama 3.1 8B table, the 3.42 GiB IQ3_S build generated 69.31 tokens per second and the larger 4.36 GiB Q4_K_S build generated 76.71. Check that both builds ran fully on the GPU before you blame the quant.

How do I know if my tokens per second is normal for my card?

Find reports with the same model build, quantization and engine on a single card like yours. Then check your model is fully on the GPU. In Ollama, ollama ps should say 100% GPU. If both match and you're still far off, look at drivers, power limits and background load.

Is tokens per second a useful metric at all?

Yes, as long as you know which one you're reading. Output speed tells you how fast text streams once it starts. Time to first token tells you how long you stare at a blank screen, and long prompts make that the bigger number.

Why does the same prompt come back so much faster the second time?

The model is already loaded, and some engines reuse the work from an identical prompt. Ollama keeps a model in memory for 5 minutes by default, so a repeat inside that window skips loading entirely.

Does a server's tokens per second tell me what one user will get?

No. A server figure often adds up the output of every active request. Look for a per-request figure or time to first token measured with one user.

Where can I compare my speed with other people on the same GPU?

The llamaperf GPU pages list community reports per card with the model, quantization and engine each one used. Open the original post before you compare, since settings that aren't recorded can still differ.