Measurement
How to benchmark Ollama generation speed with its API
By llamaperf · · 5 min read
Quick answer
For a finished Ollama request, divide eval_count by eval_duration converted to seconds. That's your output speed. Keep it apart from total request time, and write down the model, prompt, context, hardware and whether the model was already loaded.
Save a complete response
Ollama's generate endpoint returns one JSON object when you turn streaming off, which makes a run easy to keep.
Send one non-streaming request
Swap the placeholder for the exact name of a model you've installed (ollama list shows them). Keep the prompt and options identical across every run you plan to compare.
curl http://localhost:11434/api/generate \
-H 'Content-Type: application/json' \
-d '{
"model": "YOUR_INSTALLED_MODEL",
"prompt": "Explain how a hash table handles collisions, with an example.",
"stream": false,
"options": {"num_predict": 256}
}' > ollama-run.json
For a quick look without saving anything, ollama run with --verbose prints the same timings after each reply, including an eval rate.
Check the response before you do any maths
num_predict caps the output. The model can stop sooner, so read the real token count from the response instead of assuming 256. Look for an error in the file too. A failed request tells you nothing about speed, and it shouldn't end up in your spreadsheet as a zero.
Work out output speed
Two fields give you the output speed, and the only trap is the unit.
The formula
The Ollama API reference defines eval_count as the number of output tokens and describes eval_duration like this:
"Time spent generating tokens in nanoseconds" (Ollama API reference, eval_duration)
Convert to seconds first:
output tokens/second = eval_count / (eval_duration / 1,000,000,000)
The sample response on that page shows why the unit matters. It has an eval_count of 18 and an eval_duration of 52,479,709, which is about 0.05 seconds. Forget to convert and you'll publish a speed that's off by a factor of a billion.
A worked example
Say a run produced 200 tokens and eval_duration came to 10 seconds. That's 20 tokens per second. (Those are round numbers picked to show the arithmetic.)
Keep the JSON file next to your result. With only the final figure saved, you can't go back and catch a unit slip, or notice that one run stopped after twelve tokens.
Keep total time separate
The same response reports total_duration, load_duration and prompt_eval_duration as well. Each one means something different, and mixing them up is the most common way to get a wrong speed.
Loading can take longer than the answer
In the sample response on Ollama's API page, load_duration is 101,397,084 nanoseconds out of a total_duration of 174,560,334, so loading took about 58% of the request. Divide output tokens by total time on a cold start and you'll blame the GPU for time spent reading the file from disk.
How long a model stays warm is set in the Ollama FAQ:
"By default models are kept in memory for 5 minutes before being unloaded." (Ollama FAQ)
So a test you run six minutes after the last one is a cold start again.
Time to first token needs streaming
This non-streaming request can't tell you how long it took for the first word to show up. You need a streaming request for that, plus a clear rule for what counts as the first output. Write down how you measured it.
For choosing what to use every day, also time the whole task. A model that generates quickly but writes three times as much, or needs a follow-up correction, can still take longer to get you a usable answer.
Repeat it the same way every time
One run is an anecdote. Five runs with the same settings start to look like a measurement.
Write the protocol down
Decide up front whether you're timing a cold start or a model that's already loaded, and keep those in separate groups. If you send the same prompt again and again, note it, because caching can make later runs do less work.
Record the model name, Ollama version, operating system, hardware, context setting and anything else running on the machine. When you publish a summary, say what it is (the median of five runs, for example) and include the individual values. Our report template has a field for each.
Watch for answers that end early
If an answer ends almost immediately, read it before it goes into a comparison. A refusal or an early stop still produces a timing, and that timing measures the wrong work.
If your number differs from someone else's on the same card, why the same GPU can show very different speeds lists what to line up first. The benchmark reading guide covers what to record, and the submission form is where to share it. A result with clear settings is useful even when it's slower than someone else's.
Frequently asked questions
How do I see my tokens per second in Ollama?
Run ollama run with the --verbose flag and it prints the eval rate after each reply. Through the API, divide eval_count by eval_duration after converting nanoseconds to seconds.
What's the difference between prompt eval rate and eval rate?
Prompt eval rate is how fast Ollama reads your input. Eval rate is how fast it writes the answer. Reading is usually much faster, so the two numbers shouldn't be compared with each other.
Why is my first request so much slower than the rest?
The first request loads the model from disk. Ollama then keeps it in memory for 5 minutes by default, so requests inside that window skip the load.
Should I divide by total duration or by eval duration?
Use eval_duration for output speed. Total duration includes loading and prompt processing, which makes a cold start look like a slow GPU.
Does num_predict make the model write exactly that many tokens?
No. It's a cap. The model can stop earlier, so always read the real count from eval_count.
Should I report the median or the average of my runs?
The median is harder for one odd run to drag around, so it's the better headline. Publish the individual values next to it so others can see the spread.