Hardware planning
Sharing one local LLM with a few people
By llamaperf · · 5 min read
Quick answer
One loaded model can answer several people at once through parallel slots. Each slot needs its own context memory, so four slots with an 8K context need room for 32K tokens of cache. The total output goes up with more users, but each person's answers come out slower. Measure with the number of people you expect, all asking at once.
How parallel slots work
A local server keeps one copy of the model in memory and gives each request its own slot, with its own conversation state.
Ollama processes one request at a time by default
Ollama controls this with an environment variable. From its FAQ:
The maximum number of parallel requests each model will process at the same time, default 1.
From the Ollama FAQ
With the default, a second person's request waits in line until the first answer finishes. Raise OLLAMA_NUM_PARALLEL and Ollama handles that many requests together.
llama.cpp calls them slots
llama-server sets the number with -np (or --parallel), and the server README describes it as the number of server slots. Each slot is one conversation the server can work on at the same time. Pick a number close to how many people will really type at once. That's usually fewer than the number of people who have access.
Every slot needs its own context
Everyone shares the weights. Each slot keeps its own cache of the tokens it has seen, and that cache takes memory.
Multiply the context by the slot count
Ollama's FAQ says it directly: required RAM scales by OLLAMA_NUM_PARALLEL times OLLAMA_CONTEXT_LENGTH. So if one person with an 8K context fits comfortably, four slots at 8K need cache room for 32K tokens on top of the weights. That's often the thing that pushes a setup past the card's memory. Our post on context length and KV cache memory explains why the cache grows the way it does, and the calculator shows the cache size for a given model and context.
Ways to fit more people
You can give each slot a shorter context if your users send short questions. Quantizing the KV cache saves memory too, and the Ollama FAQ describes a q8_0 cache as using about half the memory of the default. A smaller model file frees room as well. Try one change at a time and check answers still come back correct.
Per-user speed and total throughput
These are the two numbers people mix up most when they talk about serving.
The total rises while each reply slows
When four people ask at once, the server does more work per second in total, because it reads the weights once for several requests. Each individual answer still comes out more slowly than it would for one person alone. A server log that reports a big combined figure tells you nothing about how the chat feels to someone waiting.
Test with real concurrency
Send the number of requests you expect at the same time, from a small script or a few browser tabs, with realistic prompt lengths. Record each person's wait before the first word and their output speed, plus the total. Run it again with one user so you have a baseline to compare. When you share the result, say how many requests ran together. Our performance guide explains why a multi-user total and a single-user speed belong in separate columns.
Plan for the busy moments
Sooner or later everyone asks at once. Decide what should happen then.
What happens when the slots are full
Extra requests wait in a queue. Ollama's FAQ says it queues up to 512 by default before it starts refusing, and that an overloaded server answers with a 503 error. For a handful of people you'll rarely hit that, but a slow model with long answers can make the wait feel long well before the limit. If waits get annoying, fewer slots with shorter contexts sometimes feel faster than many slow ones.
When a serving engine makes sense
For a small group, Ollama or llama-server is enough. If you're serving a team through an API all day, look at an engine built for that. The vLLM README lists continuous batching of incoming requests among its features, which keeps the GPU busy as requests come and go. Browse GPU reports for multi-user results before you buy hardware for it, and submit your own with the concurrency noted.
Frequently asked questions
Can a single loaded model handle parallel requests?
Yes. Ollama and llama-server both keep one copy of the model and serve several requests at once through parallel slots. Ollama processes one request at a time unless you raise OLLAMA_NUM_PARALLEL.
How many concurrent users can a local LLM handle?
It depends on your memory and how long everyone's context is. Each extra slot needs its own cache, so work out the memory for one user and multiply. Then test with that many people asking at once.
How do parallel requests share context size in llama.cpp?
llama-server can keep one KV cache buffer that every slot draws from. Its README says this unified buffer is on by default when the slot count is automatic, and the kv-unified-per-slot option caps how much any one slot can use.
Does OLLAMA_NUM_PARALLEL use more memory?
Yes. The Ollama FAQ says required RAM scales by OLLAMA_NUM_PARALLEL times OLLAMA_CONTEXT_LENGTH, so doubling the slots roughly doubles the memory the context needs.
Why is each user slower when several people use the model?
The GPU splits its work between the requests. Total output goes up, but each person's share of it goes down, so individual answers take longer than they would for one user alone.
Should I use vLLM instead of Ollama for multiple users?
For a few people, Ollama or llama-server is fine. For a team hitting an API all day, an engine with continuous batching such as vLLM is built for that load. Test both with your real traffic if you're unsure.