llamaperf

The llamaperf blog

Practical ways to choose a local LLM, measure its speed and improve your setup.

Measurement · · 5 min read

How to measure tokens per watt on your local LLM

Log your GPU's power draw with nvidia-smi, line it up with a timed run and work out the energy each answer costs. Here's the method and the traps to avoid.

Hardware planning · · 5 min read

Sharing one local LLM with a few people

Serving a local model to your team or family? Here's how parallel slots work, why each one costs memory, and how to read per-user speed against the total.

Hardware planning · · 6 min read

Running a local LLM on CPU only: what to expect

No graphics card? You can still run a local model. Here's why memory bandwidth sets your speed, and how to test a CPU-only setup before you buy anything.

Choosing a model · · 6 min read

MoE models: total vs active parameters

A mixture-of-experts model uses only a few experts per token, yet all of them have to sit in memory. Here's what that means for your hardware.

Hardware planning · · 5 min read

Why a longer context needs more memory

Every token in your context keeps its keys and values in memory. Here's how that cache grows, how to size it and how to shrink it in Ollama.