Troubleshooting
When a model doesn't fit in VRAM: partial offload
By llamaperf · · 5 min read
Quick answer
The runtime keeps as many layers as fit on the GPU and runs the rest on the CPU from system RAM. It works, and it's a lot slower, because every token still passes through the CPU layers. Check the split with ollama ps or the llama.cpp startup log, then compare against a smaller quant that fits completely.
What happens when it spills
You download a model that's a little too big for your card, and it still loads. That's partial offload doing its job. llama.cpp lists it as a core feature:
CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity From the llama.cpp README
Layers split between GPU and CPU
A model is a stack of layers, and each token goes through all of them in order. With partial offload, the first chunk of layers lives in VRAM and runs on the GPU, and the rest lives in system RAM and runs on the CPU. Ollama and LM Studio build on llama.cpp and make the same split for you when the model doesn't fit.
Why a small spill hurts so much
Generating a token means reading every weight once. System RAM is much slower to read than GPU memory, so the CPU layers take far longer per byte than the GPU layers. The two parts run one after the other, so the times add up.
Here's an example with round numbers. Say the GPU part of a token takes 10 ms and the CPU part takes another 10 ms. You've gone from 100 tokens per second to 50, even though most of the model is still on the GPU. That's why people are surprised when moving just a few layers to RAM cuts their speed in half.
Check where the layers went
Before you change any settings, find out what split you're running. A lot of "my GPU is slow" posts turn out to be a model that quietly ended up partly on the CPU.
In Ollama
Run ollama ps while the model is loaded. The Processor column tells you the split. Ollama's FAQ describes the mixed case like this:
48%/52% CPU/GPU means the model was loaded partially onto both the GPU and into system memory From the Ollama FAQ
100% GPU is what you want to see. Anything else means some of the model is in system RAM. Save that line alongside any speed you record.
In llama.cpp
The startup log prints a line such as offloaded 30/33 layers to GPU. If it says 0/33, nothing reached the GPU. Usually that means your build has no GPU backend or the layer count is set to zero. Check that the log names your GPU before you blame the model.
Setting the layer count
If you run llama.cpp directly, you choose how many layers go to the GPU.
The -ngl option
The server documentation describes -ngl (also --n-gpu-layers) as the "max. number of layers to store in VRAM". Set it to a number, or to all to try to put everything on the GPU. If loading fails with an out-of-memory error, lower it a few layers at a time until it loads, and note the highest value that works.
Leave room for the context
The layers you move to the GPU compete with the KV cache for the same VRAM. A layer count that loads fine at 4K can run out of memory at 32K. Our post on context length and the KV cache explains why. Set the context you need first, then find the layer count, and test with a long prompt as well as a short one. Keep some VRAM free for your desktop too, because the display uses the same card.
Offload or a smaller quant?
Partial offload keeps a bigger model within reach. A smaller quant that fits completely is usually much faster. Which one is better depends on whether the bigger model gets your work right.
Run the same prompt both ways
Take one of your real tasks and run it twice: the bigger quant with partial offload, and a smaller quant that shows 100% GPU. Our guide to GGUF quant names helps you pick the smaller file, and the VRAM calculator shows whether it fits with your context. Record the time to a finished answer and whether the answer passed. If both pass, the faster one wins. If only the offloaded model gets it right, the wait may be worth it for that job.
Write down the split
When you compare numbers with other people, the split matters as much as the card. Two people with the same GPU can report very different speeds because one of them had six layers on the CPU. The benchmark-reading guide lists what to record. If you submit a report, include the layer split or the ollama ps line, so your number lands next to comparable ones on the GPU pages.
Frequently asked questions
How many layers can I offload to my GPU?
As many as fit alongside the context cache and your desktop's own use of the card. Start with all of them, and if loading fails, lower the count a few layers at a time. The highest value that loads with your real context is the one to keep.
Is the speed loss from moving layers to the CPU linear?
No. The CPU layers read from slower system RAM, so each one costs much more time than a GPU layer. A few layers on the CPU can take as long as all the GPU layers together, which is why speed can halve even when most of the model is on the GPU.
Why does llama.cpp say offloaded 0 layers to GPU?
Nothing went to the GPU. Usually the build has no GPU backend, or the layer count is set to zero. Check that the startup log names your GPU, and install a build with CUDA, ROCm, Vulkan or Metal support for your hardware.
How can I tell if Ollama is running my model on the GPU?
Run ollama ps while the model is loaded. The Processor column shows 100% GPU when everything is in VRAM, 100% CPU when it's all in system memory, and a split like 48%/52% CPU/GPU when it's divided between them.
Is it better to offload a bigger model or run a smaller quant fully on the GPU?
Test both on a real task. The smaller quant is usually much faster. Keep the offloaded model only if it gets answers right that the smaller one misses and the extra wait doesn't bother you.