Choosing a model
Ollama, llama.cpp, vLLM or MLX-LM: picking an engine
By llamaperf · · 5 min read
Quick answer
Pick the engine by how you'll use the model. For one person on a PC or Mac, Ollama or llama.cpp is the easy start. For serving many requests from NVIDIA or AMD GPUs, look at vLLM. On Apple silicon, MLX-LM is worth trying next to llama.cpp. Then run the same model and prompt on two of them and keep the one that fits your routine.
Start from how you'll use it
The right engine depends far more on your situation than on any ranking. Answer one question first: who is going to send it requests?
Just you, on your own machine
If you're the only user, you want quick setup and good speed on consumer hardware. Ollama and llama.cpp both fit. Ollama wraps a model library and a simple API around the engine, and llama.cpp gives you the engine directly with every setting exposed. If you're on a Mac, MLX-LM is also in the running.
Other people or programs, all day
If a team or an app will call the model constantly, you care about how many requests it can handle together without each one crawling. That's what serving engines like vLLM are designed for. Our post on sharing one model with a few people covers where a small group stops needing one.
What each engine is built for
Each project describes its own goal clearly, and the descriptions tell you a lot.
llama.cpp and Ollama
llama.cpp runs on almost anything, from a laptop CPU to a multi-GPU box, and its README opens its feature list with this:
Plain C/C++ implementation without any dependencies
From the llama.cpp README
That's why so many other tools use it underneath. Ollama uses the same GGUF format and is one of the easiest ways to get a model running. You pull a model by name, and the Ollama docs on importing models show how to bring in your own GGUF or safetensors files when the library doesn't have what you want.
vLLM and MLX-LM
vLLM targets serving. Its README describes it this way:
vLLM is a fast and easy-to-use library for LLM inference and serving.
From the vLLM README
The same README lists continuous batching and an OpenAI-compatible API server among its features. MLX-LM is the MLX project's package for running and fine-tuning models on Apple silicon, and it pulls ready-converted models from the MLX Community organisation on Hugging Face.
Model formats decide a lot
Your choice of engine also picks the file formats you can use, and the format decides how much memory the model needs.
GGUF and its quantizations
llama.cpp and Ollama use GGUF files, which come in many quantization levels. The llama.cpp quantize README lists Llama 3.1 8B at 14.96 GiB in F16 and 4.58 GiB at Q4_K_M, less than a third of the size. That's what makes big models fit on ordinary hardware. Our post on what Q4_K_M means explains the names, and the quantization guide covers the trade-offs.
Check what your hardware runs
Serving engines usually expect the model in its Hugging Face form or in a quantization built for GPUs, so the memory you need can be very different for the same model. On a Mac, MLX-LM uses MLX-converted weights while llama.cpp uses GGUF, and both use Apple's GPU. Check the calculator for the memory a given quantization needs on your card, and our Mac pages for what people run on each chip.
Try two before you commit
Reading about engines only gets you so far. An afternoon of testing settles it.
Use the same model and prompt
Pick one model and get it in the format each engine uses, at similar quantization. Keep the prompt and context size identical, cap the output at the same length, and time the wait before the first word and the total answer time on both. Different engines report speed in slightly different ways, so time things the same way yourself where you can.
Keep the one that fits your routine
Speed matters, but so does how you'll live with it. Updates, model downloads, the API your other tools expect and how easy it is to change settings all come into it. Browse the engine pages to see what other people run on similar hardware. When you've made your choice, submit a report with the engine and version noted, so the next person picking has one more data point.
Frequently asked questions
Should I use Ollama or llama.cpp?
Ollama is quicker to set up and manages models for you. llama.cpp exposes every setting and gets new features first. Both run the same GGUF files, so you can start with one and switch later.
When should I use vLLM instead of Ollama?
When many requests hit the model at the same time, such as a team or an app calling an API all day. vLLM is built for serving with continuous batching. For one person, Ollama or llama.cpp is simpler.
Is MLX faster than llama.cpp on a Mac?
It depends on the model and the quantization. Both use the Mac's GPU. Run the same model in each format with the same prompt and compare the times on your own machine.
Can Ollama run GGUF files from Hugging Face?
Yes. The Ollama docs describe importing a GGUF file through a Modelfile that points at it, and the same page covers importing safetensors weights.
Does vLLM use less VRAM than llama.cpp?
Not as a rule. The memory you need depends mainly on the model format and quantization each engine loads, and on the context and number of parallel requests. Compare at matched settings.
Can I switch engines without downloading the model again?
Only if the new engine reads the same format. llama.cpp and Ollama both use GGUF, while vLLM and MLX-LM usually need a different download or a conversion.