llamaperf
← All articles

Measurement

How to measure time to first token on a local LLM

By llamaperf · · 6 min read

Quick answer

Send a streaming request, start a clock when you send it, and stop it when the first chunk with visible text arrives. Write down the prompt length next to the result. A cold start or a cached prompt can change the number by seconds, so note those too.

What the wait is made of

Time to first token, or TTFT, is how long you stare at an empty reply before text appears. For chat it's often what makes a setup feel slow, even when the words come quickly once they start.

What happens before the first word

Three things happen before the first token. If the model isn't in memory yet, the runtime loads it from disk. Then it reads your whole prompt, which is called prompt processing or prefill. Then it generates the first token. Ollama reports these separately in its API response: load_duration is "Time spent loading the model in nanoseconds" and prompt_eval_duration is "Time spent evaluating uncached prompt tokens in nanoseconds".

Why long prompts wait longer

Prompt processing has to get through every token you send, including the whole chat history. A 20,000-token document takes far longer to read than a one-line question, even on the same setup. So a TTFT figure means nothing without the prompt length next to it. If you want the background on why long inputs also need more memory, see our post on context length and the KV cache.

Timing it with Ollama

The simplest reliable way is a small script that sends a streaming request and notes when the first piece of text arrives. Streaming is Ollama's default:

When true, returns a stream of partial responses From the stream parameter in the Ollama generate API reference

A small streaming script

This uses only Python's standard library. Put your prompt in prompt.txt and swap in a model you have installed.

import json, time, urllib.request

body = json.dumps({
    "model": "YOUR_INSTALLED_MODEL",
    "prompt": open("prompt.txt").read(),
    "stream": True,
}).encode()
req = urllib.request.Request(
    "http://localhost:11434/api/generate", data=body,
    headers={"Content-Type": "application/json"})

start = time.perf_counter()
first = None
with urllib.request.urlopen(req) as resp:
    for line in resp:
        chunk = json.loads(line)
        if first is None and (chunk.get("response") or chunk.get("thinking")):
            first = time.perf_counter() - start
        if chunk.get("done"):
            final = chunk

print(f"time to first token: {first:.2f} s")
print("load s:", final["load_duration"] / 1e9)
print("prompt tokens:", final["prompt_eval_count"])
print("cached prompt tokens:", final.get("prompt_eval_cached_count"))
print("prompt s:", final["prompt_eval_duration"] / 1e9)

What counts as the first token

Reasoning models stream their thinking before the answer. The script above counts the first thinking text as the first token. If you care about when the answer itself starts, check only response instead. Either choice is fine, as long as you say which one you used and stick with it.

Timing it with llama.cpp's server

llama.cpp's server gives you most of the breakdown without a script, and the same streaming approach works if you want a wall-clock figure.

The timings object

Every response from the server includes a timings object. The server README documents prompt_n (prompt tokens processed), cache_n (prompt tokens reused from the cache) and prompt_ms, the time spent on the prompt. prompt_ms plus the time for one generated token is a close estimate of TTFT from the server's side. A client-side stopwatch adds network and client time on top, so don't mix the two in one comparison.

Prompt caching

The server reuses work from the previous request when the start of the prompt matches:

Re-use KV cache from a previous request if possible. This way the common prefix does not have to be re-processed, only the suffix that differs between the requests. From the cache_prompt option in the llama.cpp server README

That helps in daily use. In a benchmark it can fool you. Send the same prompt twice and the second TTFT can drop to a fraction of the first, because almost nothing was processed. Check cache_n, or Ollama's cached prompt count, before you believe a fast result. It also explains why a chat tool that rewrites the start of its system prompt on every turn feels slow: the cache can't match, so everything is read again.

A test you can repeat

A single TTFT number is easy to misread. A small, fixed routine makes it worth sharing.

Separate cold starts from warm runs

Measure each case on its own. Cold means the model wasn't loaded, and load_duration will be large. Warm means the model is loaded and the prompt is new. Cached means you repeated a prompt that was already processed. For comparing hardware or settings, warm runs with fresh prompts are the ones to use. Change a word at the start of the prompt between runs to defeat the cache.

What to record

FieldWhy
Prompt tokensTTFT grows with prompt length
Cached prompt tokensA cache hit hides prompt processing
Load timeSeparates a cold start from a warm run
First-token ruleThinking text or answer text
Timing sourceServer timings or client stopwatch
RunsReport the median and keep every value

Try a short, a typical and a long prompt from your real work, so you can see how the wait grows. The performance guide explains how TTFT relates to generation speed, and the benchmark-reading guide covers the rest of a clean report. Our report template has fields for both numbers. When you've got results you trust, submit them with the prompt length included.

Frequently asked questions

How do I measure time to first token in Ollama?

Send a streaming request to the generate endpoint, start a timer when you send it, and stop it when the first chunk with response or thinking text arrives. Record load_duration and prompt_eval_count from the final chunk so you know what the wait included.

Why is the first response slow and later ones fast?

The first request may include loading the model from disk, and repeat prompts can reuse cached work. Check load_duration and the cached prompt token count. A warm run with a fresh prompt is the fair comparison.

Is time to first token the same as prompt processing time?

Nearly, on a warm model. TTFT is prompt processing plus generating the first token, plus loading if the model wasn't in memory. A client-side stopwatch also adds a little network and client time.

How do I make prompt processing faster?

Send fewer tokens and keep the start of your prompt stable so the cache can reuse it. Keeping the whole model on the GPU helps too, because prompt processing on the CPU is far slower.

Why does my coding agent reprocess the whole prompt every turn?

The cache only reuses a prefix that matches exactly. If the tool changes something near the start of the prompt, such as a timestamp or mode switch in the system prompt, everything after it has to be read again.

What's an acceptable time to first token?

It depends on the job. For chat, most people want text within a couple of seconds. For a long document, a longer wait is normal. Measure your own prompts and decide what you'll put up with.