Choosing a model
How to choose a local LLM for the hardware you already own
By llamaperf · · 6 min read
Updated
Quick answer
Pick the model that gets your real tasks right at the context you need, and answers fast enough that you'll keep using it. Start from the machine you have. Shortlist two setups in the calculator and run the same prompts on both before you download anything bigger.
Start with the job you need done
Before you open a leaderboard, write down what you want the model to do. "Help with code" is too vague to test. "Fix this failing function without changing its signature" gives you something you can check in a minute.
Build a small prompt set
Collect five or six prompts from your real week. Put in something easy you'd ask every day and something hard you've struggled with. Add one long input too, since that's where memory runs out first. Decide what a pass looks like for each prompt before you run anything.
For document work, use a file where you already know which details matter. You'll spot a missed detail in seconds. Nobody else's prompts can tell you which model suits your work.
Score the result apart from the writing
A polished explanation that ignores the output format you asked for is a miss. Keep two columns in your notes. One says whether the model did the task. The other says whether you liked reading the answer. The first column decides.
Size the setup for the context you'll use
Open the VRAM calculator and pick your hardware. Set the context to what you'll really send. Short chat turns need far less than whole reports pasted in, and planning for the biggest window on offer rules out models that would've served you fine.
Check your runtime's default
Your runtime picks a context size too, and it may not be the one you expect. Ollama's context length page says it defaults to 4k tokens on machines with under 24 GiB of VRAM, 32k from 24 to 48 GiB, and 256k at 48 GiB and above. The same page is blunt about the cost:
"Setting a larger context length will increase the amount of memory required to run a model." (Ollama, Context length)
So look at the whole memory breakdown in the calculator, including the space inference needs on top of the weights.
Leave room for the rest of your machine
The calculator gives you an estimate. Your machine has the final say, because other apps and the exact model file both move the numbers. Keep some headroom so the computer stays usable, then check what your inference tool reports once the model loads.
On a Mac, macOS keeps part of the unified memory for itself, so a 32 GB Mac doesn't hand the model 32 GB. The Mac pages group community reports by chip and memory size, and the calculator shows how much of that memory it assumes the GPU can use.
Shortlist exact builds
Pick one small, quick option and one bigger option that still fits. For each, save the exact model name, the quantization, the runtime version and the context setting. A family name on its own doesn't tell anyone what you ran.
How much the quantization changes the size
File size swings a lot with quantization. The llama.cpp quantize README lists Llama 3.1 8B at 32.1 GB in its original form and 4.9 GB at Q4_K_M, and the 70B at 280.9 GB against 43.1 GB. So "can I run the 70B?" is a question about which build of it you mean. Our quantization guide explains what the labels stand for.
Read reports from machines like yours
Browse hardware reports for setups close to yours, and open the original post when there's one. A report with a different quantization or two GPUs still gives you background. It won't predict your speed.
When nobody has reported your combination, write "unknown" in the speed column and go measure it. That's more honest than borrowing the nearest number, and it's a gap worth filling.
Choose on work that got done
Run your prompts through both options and note whether each one passed. Then time two things: the wait before text appears, and how long the useful answer took to finish. Run everything at least twice so one lucky fast run doesn't decide it.
Keep the results in one table
| Question | What to record |
|---|---|
| Did it solve the task? | Pass criteria and any manual repair |
| Was the wait acceptable? | Initial wait and full completion time |
| Did it hold up on longer input? | Context setting and input length |
| Can you reproduce it? | Model, quantization, engine and hardware |
The benchmark reading guide covers what else is worth writing down.
When the bigger model earns its place
If the smaller model passes as often, keep it. You'll get faster answers and more free memory. Keep the bigger one when it wins tasks the small one keeps failing and the wait doesn't bother you.
When your work changes, rerun the same prompts so you can see whether a new setup helped. And if you land on a combination nobody has posted yet, submit it. The next person with your machine gets a real number to start from.
Frequently asked questions
Is there a way to estimate tokens per second from my VRAM?
VRAM tells you what fits. Speed depends mostly on how fast the card can read the model's weights from memory, so two cards with the same VRAM can differ a lot. The llamaperf calculator estimates speed from the card's memory bandwidth, and a measured report on the same card and build is better than any estimate.
What's the minimum tokens per second that's still usable?
There's no single floor. Reading along in a chat feels fine at a much lower speed than an agent that has to read and act on every answer. Time the wait on your own prompts and ask whether you'd put up with it every day.
How much memory does a model need?
Start from the file size of the build you download, then add the context cache and some runtime overhead. For scale, the llama.cpp quantize README lists Llama 3.1 8B at 4.9 GB in Q4_K_M and the 70B at 43.1 GB.
Should I run a bigger model at a lower quant or a smaller model at a higher quant?
Test both on your own prompts. The calculator shows which of them fit your memory, and your pass rate on real tasks tells you which one to keep.
Does the context length setting matter if I only ask unrelated questions?
Within one chat, earlier turns are sent again with every new message, so they count toward the context. A fresh chat starts empty. The setting itself still decides how much memory the runtime may use, so size it for your longest real conversation.