Measurement
How to measure time to first token on a local LLM
By llamaperf · · 6 min read
Quick answer
Send a streaming request, start a clock when you send it, and stop it when the first chunk with visible text arrives. Write down the prompt length next to the result. A cold start or a cached prompt can change the number by seconds, so note those too.
What the wait is made of
Time to first token, or TTFT, is how long you stare at an empty reply before text appears. For chat it's often what makes a setup feel slow, even when the words come quickly once they start.
What happens before the first word
Three things happen before the first token. If the model isn't in memory yet, the runtime loads it from disk. Then it reads your whole prompt, which is called prompt processing or prefill. Then it generates the first token. Ollama reports these separately in its API response: load_duration is "Time spent loading the model in nanoseconds" and prompt_eval_duration is "Time spent evaluating uncached prompt tokens in nanoseconds".
Why long prompts wait longer
Prompt processing has to get through every token you send, including the whole chat history. A 20,000-token document takes far longer to read than a one-line question, even on the same setup. So a TTFT figure means nothing without the prompt length next to it. If you want the background on why long inputs also need more memory, see our post on context length and the KV cache.
Timing it with Ollama
The simplest reliable way is a small script that sends a streaming request and notes when the first piece of text arrives. Streaming is Ollama's default:
When true, returns a stream of partial responses From the
streamparameter in the Ollama generate API reference
A small streaming script
This uses only Python's standard library. Put your prompt in prompt.txt and swap in a model you have installed.
import json, time, urllib.request
body = json.dumps({
"model": "YOUR_INSTALLED_MODEL",
"prompt": open("prompt.txt").read(),
"stream": True,
}).encode()
req = urllib.request.Request(
"http://localhost:11434/api/generate", data=body,
headers={"Content-Type": "application/json"})
start = time.perf_counter()
first = None
with urllib.request.urlopen(req) as resp:
for line in resp:
chunk = json.loads(line)
if first is None and (chunk.get("response") or chunk.get("thinking")):
first = time.perf_counter() - start
if chunk.get("done"):
final = chunk
print(f"time to first token: {first:.2f} s")
print("load s:", final["load_duration"] / 1e9)
print("prompt tokens:", final["prompt_eval_count"])
print("cached prompt tokens:", final.get("prompt_eval_cached_count"))
print("prompt s:", final["prompt_eval_duration"] / 1e9)
What counts as the first token
Reasoning models stream their thinking before the answer. The script above counts the first thinking text as the first token. If you care about when the answer itself starts, check only response instead. Either choice is fine, as long as you say which one you used and stick with it.
Timing it with llama.cpp's server
llama.cpp's server gives you most of the breakdown without a script, and the same streaming approach works if you want a wall-clock figure.
The timings object
Every response from the server includes a timings object. The server README documents prompt_n (prompt tokens processed), cache_n (prompt tokens reused from the cache) and prompt_ms, the time spent on the prompt. prompt_ms plus the time for one generated token is a close estimate of TTFT from the server's side. A client-side stopwatch adds network and client time on top, so don't mix the two in one comparison.
Prompt caching
The server reuses work from the previous request when the start of the prompt matches:
Re-use KV cache from a previous request if possible. This way the common prefix does not have to be re-processed, only the suffix that differs between the requests. From the
cache_promptoption in the llama.cpp server README
That helps in daily use. In a benchmark it can fool you. Send the same prompt twice and the second TTFT can drop to a fraction of the first, because almost nothing was processed. Check cache_n, or Ollama's cached prompt count, before you believe a fast result. It also explains why a chat tool that rewrites the start of its system prompt on every turn feels slow: the cache can't match, so everything is read again.
A test you can repeat
A single TTFT number is easy to misread. A small, fixed routine makes it worth sharing.
Separate cold starts from warm runs
Measure each case on its own. Cold means the model wasn't loaded, and load_duration will be large. Warm means the model is loaded and the prompt is new. Cached means you repeated a prompt that was already processed. For comparing hardware or settings, warm runs with fresh prompts are the ones to use. Change a word at the start of the prompt between runs to defeat the cache.
What to record
| Field | Why |
|---|---|
| Prompt tokens | TTFT grows with prompt length |
| Cached prompt tokens | A cache hit hides prompt processing |
| Load time | Separates a cold start from a warm run |
| First-token rule | Thinking text or answer text |
| Timing source | Server timings or client stopwatch |
| Runs | Report the median and keep every value |
Try a short, a typical and a long prompt from your real work, so you can see how the wait grows. The performance guide explains how TTFT relates to generation speed, and the benchmark-reading guide covers the rest of a clean report. Our report template has fields for both numbers. When you've got results you trust, submit them with the prompt length included.
Frequently asked questions
How do I measure time to first token in Ollama?
Send a streaming request to the generate endpoint, start a timer when you send it, and stop it when the first chunk with response or thinking text arrives. Record load_duration and prompt_eval_count from the final chunk so you know what the wait included.
Why is the first response slow and later ones fast?
The first request may include loading the model from disk, and repeat prompts can reuse cached work. Check load_duration and the cached prompt token count. A warm run with a fresh prompt is the fair comparison.
Is time to first token the same as prompt processing time?
Nearly, on a warm model. TTFT is prompt processing plus generating the first token, plus loading if the model wasn't in memory. A client-side stopwatch also adds a little network and client time.
How do I make prompt processing faster?
Send fewer tokens and keep the start of your prompt stable so the cache can reuse it. Keeping the whole model on the GPU helps too, because prompt processing on the CPU is far slower.
Why does my coding agent reprocess the whole prompt every turn?
The cache only reuses a prefix that matches exactly. If the tool changes something near the start of the prompt, such as a timestamp or mode switch in the system prompt, everything after it has to be read again.
What's an acceptable time to first token?
It depends on the job. For chat, most people want text within a couple of seconds. For a long document, a longer wait is normal. Measure your own prompts and decide what you'll put up with.