llamaperf
← All articles

Measurement

How to measure tokens per watt on your local LLM

By llamaperf · · 5 min read

Quick answer

Log your GPU's power once a second with nvidia-smi while a timed run is going. Multiply the average watts by the run time to get joules, then divide by the tokens produced. Compare setups on energy per token with the same prompt and output length, and remember the GPU reading leaves out the rest of the computer.

Why energy per answer is the number you want

A card's wattage on its own says little. A fast card that draws more power can still use less energy per answer, because it finishes sooner.

Watts are a rate, your bill counts energy

Power is how fast energy is used. Energy is power times time. So the useful figure for a local model is energy per token (joules per token), or its inverse, tokens per joule, which is what people usually mean by tokens per watt. It rewards a setup that gets the work done, whether it draws a lot for a short time or a little for longer.

Idle draw counts too

A card that sits loaded and waiting all day still uses power between questions. If your model runs around the clock, note the idle reading as well as the reading under load. It can matter more than the load figure for a machine that answers a few questions an hour.

Log power with nvidia-smi

On an NVIDIA card, nvidia-smi can print power readings on a loop and write them to a file.

The command

nvidia-smi --query-gpu=timestamp,power.draw.average --format=csv -l 1 > power.csv

The nvidia-smi documentation defines the field this reads:

The average power draw for the entire board for the last second, in watts.

From the nvidia-smi documentation

The same page says that field is supported on Ampere (except GA100) and newer cards. On an older card, query power.draw instead, which is the last measured reading. The -l 1 sets a one-second interval. Leave it out and you get a single reading, and the docs note that -l with no number defaults to every 5 seconds, which is too coarse for short runs.

Line the log up with the run

Start the log, wait a few seconds for an idle baseline, run your prompt, then wait a few seconds more. Keep the timestamps from the log and from your run so you can cut out exactly the seconds the model was working. A timed run with a token count is what you need on the other side. Our post on measuring time to first token shows how to get both from a streaming request.

Turn the log into joules per token

Once you have power readings and a run you can match them to, the sum is short.

The arithmetic

Average the power readings from the seconds the model was generating, multiply by the generation time and divide by the tokens generated. The Ollama API reference gives both of the other numbers directly: eval_count is the output tokens and eval_duration is the time spent generating them, in nanoseconds.

Here's an example with made-up numbers. A run averages 300 W for 20 seconds and produces 600 tokens. That's 6,000 joules for 600 tokens, or 10 joules per token. Put the other way round, it's 0.1 tokens per joule.

What the GPU reading leaves out

nvidia-smi measures the card. Your CPU, memory, fans and power supply losses aren't in that number, and on a machine with partial CPU offload the CPU can do a lot of the work. If you want the whole machine, a plug-in wall meter gives you that, at a coarser interval. Say which one you measured when you share a result.

Compare setups fairly

Energy numbers are only comparable when the work is the same.

Hold the workload fixed

Use the same prompt and the same output length for every setup you test, and run each one a few times. Prompt processing and generation draw different power, so a run with a long prompt and a short answer gives a different energy per token than the reverse. Write down both lengths.

Power limits change the result

The nvidia-smi docs note that the power limit can be adjusted with -pl. Generation speed and power don't fall in step, so a lower limit can improve energy per token. Measure it on your own card. If you run with a limit, include it in your report. It's as much a part of the setup as the quantization. The report template has space for it, our benchmark-reading guide covers the other fields, and GPU reports show what others run so you can submit yours alongside them.

Frequently asked questions

How do I measure tokens per watt for a local LLM?

Log the GPU's power once a second while you time a run. Multiply the average watts by the generation time to get joules, then divide the tokens produced by the joules, or the joules by the tokens for energy per token.

Does nvidia-smi show the power draw of the whole computer?

No, only the card. The CPU, memory, fans and power supply losses aren't included. Use a plug-in wall meter if you want the whole machine.

Does lowering the GPU power limit slow down inference?

It can. Speed and power don't fall in step, so a lower limit may cost little speed and save more energy, or the reverse. Measure your own card at two or three limits and note the limit when you share results.

How much power does a GPU draw when a model is idle?

It depends on the card and driver. Log it with nvidia-smi while the model is loaded and nothing is running, and include that figure if your machine sits idle between questions most of the day.

Is a Mac more power efficient for local LLMs?

Measure rather than assume. On a Mac, the powermetrics man page says its power values are estimated and shouldn't be used to compare different devices, so treat Mac readings as a guide and prefer a wall meter for comparisons.

Why do my tokens per watt numbers change between runs?

Prompt length, output length, temperature and whatever else the machine is doing all shift the result. Keep the prompt and output length the same, repeat each run a few times and report the median.