Hardware planning
Running a local LLM on CPU only: what to expect
By llamaperf · · 6 min read
Quick answer
You can run a local model on a CPU alone, and small quantized models are pleasant to use that way. Your memory bandwidth decides how fast the answer comes out, because every new token reads the whole model from RAM. Test with a small quantized model first, time the prompt and the answer separately, and only then decide whether you need a GPU.
You can run a model with no GPU at all
Plenty of people run local models on a laptop or an old desktop with no graphics card. llama.cpp was written to run on ordinary processors first, and its README lists the instruction sets it uses on x86 chips:
AVX, AVX2, AVX512 and AMX support for x86 architectures
From the llama.cpp README
What you give up
Speed, mostly. A CPU can do the maths, but it reads memory far more slowly than a graphics card reads its own VRAM, and reading memory is most of what generation does. You'll also wait longer for long prompts, because processing your input is heavy arithmetic and a desktop CPU has much less of that to spare.
When CPU-only is the right call
It works well for short questions and for jobs you're happy to leave running in the background. It's also the cheapest way to find out whether local models are useful to you at all. Try it on the machine you have before you spend money on a card. If you later want to know where a GPU would change things, our performance guide explains which part of a run each piece of hardware speeds up.
Memory bandwidth sets the ceiling
When a model writes an answer, it produces one token at a time, and for each token it reads essentially all of its weights from memory. So the size of the model file and the speed of your RAM decide how fast text can appear.
Smaller files come out faster
Quantization shrinks the file, which is why it matters so much on a CPU. The llama.cpp quantize README lists Llama 3.1 8B at 32.1 GB in its original form and 4.9 GB at Q4_K_M. That's less than a sixth of the data to read for every token. Our post on what Q4_K_M means explains the labels if you're picking a file.
Here's the rough maths, with made-up numbers to show how it works. Say your RAM delivers 60 GB/s and the model file is 5 GB. At that rate you can read the file twelve times a second, so you'd top out around 12 tokens per second, and real runs land below that ceiling. Look up your own memory speed and put your own numbers in. The calculator does the same sum for the hardware it knows.
Why extra cores stop helping
Once the cores are waiting on memory, adding threads does nothing for generation. Dual-channel memory in a typical desktop is the limit you hit, and a workstation board with more memory channels raises it. Faster RAM helps too. A newer processor with the same memory setup often helps less than you'd hope.
Test it on your own machine
Start with a small model at a 4-bit quantization so the first run finishes quickly. Then move up in size once you know your baseline.
Time the prompt and the answer separately
Reading your prompt and writing the answer stress different parts of the machine, so time them apart. Long prompts are where a CPU hurts most. Our post on measuring time to first token shows how to get both numbers from a streaming request.
Try a few thread counts too. Start at the number of physical cores and go down from there, and keep whichever setting gives the fastest generation on your machine. Hyperthreads rarely help a job that's waiting on memory.
Write down what else was running
Background apps compete for the same memory bandwidth. Note your RAM size and speed, how many memory channels you have, and whether a browser full of tabs was open. Those details decide whether someone else can compare their result with yours. When you've got clean numbers, submit them. CPU-only results are rare in most collections, so yours helps.
Decide whether you need a GPU
Your test tells you whether the CPU is good enough for the work you do. A few signs say it isn't.
When a card starts paying off
If you keep waiting a minute or more before the first word on long documents, a GPU will shorten that wait more than anything else you could buy. The same goes if you need a bigger model than your RAM can feed at a readable pace. For models that almost fit on a card you already own, splitting the layers between GPU and CPU is a middle path, covered in our post on models too big for VRAM.
Sparse mixture-of-experts models are worth a look on a CPU because they read only part of their weights for each token. Our post on MoE memory and speed explains the trade. Whatever you pick, browse GPU reports from people running similar models before you buy, and compare like with like.
Frequently asked questions
Can you run a local LLM without a GPU?
Yes. llama.cpp and tools built on it run on ordinary x86 and ARM processors. Small quantized models answer at a readable pace on a modern desktop, and larger ones work too if you're willing to wait.
How many tokens per second can I get on CPU only?
It depends mostly on your memory bandwidth and the size of the model file. Divide your RAM bandwidth by the file size to get a rough ceiling, then measure, because real runs come in below it.
Does RAM speed matter more than CPU cores for local LLMs?
For writing the answer, yes. Generation waits on memory, so faster RAM and more memory channels help more than extra cores. Prompt processing is the part where more cores still help.
Is DDR5 worth it for CPU inference?
Faster memory raises the ceiling on generation speed, so DDR5 helps compared with slower DDR4 in the same number of channels. A board with more channels raises it further. Measure on your own setup before and after if you can.
What is the best small model to run without a GPU?
Pick a small model that does your task well and run it at a 4-bit quantization. Test two or three candidates on your own prompts, since the right choice depends on what you ask it to do.
How many threads should I use for llama.cpp on CPU?
Start at your physical core count and try a few lower values. Keep whichever gives the fastest generation. Going above the physical core count rarely helps because the threads end up waiting on memory.