Running a local LLM on two or more GPUs
What people run on two or more cards of the same kind, the speeds they report, and the same card on its own beside it.
Can you use two GPUs for a local LLM?
Yes. People run local models on two cards and on rigs of up to 12, and the setup reported most is 2× NVIDIA RTX 3090. The usual reason is memory: two cards hold a model one cannot. Where the same card also ran the same model alone, the multi-card median was faster in all 4 cases, though that is few pairs and depends on the engine splitting the model across cards.
The calculator shows what fits on one card first, which tells you whether you need a second.
One card or several, same model
Median tokens per second for the same model size and quant level on the same card, alone and in a multi-card setup. Plain runs only: one request, no speculative decoding, nothing in system RAM.
| Setup and model | One card | Several |
|---|---|---|
| 2× NVIDIA RTX 3090Qwen3.8 27B at 4-bit | 36.43 runs | 55.02 runs |
| 2× AMD Radeon AI PRO R9700 32GBQwen3.8 27B at 4-bit | 29.0one run | 37.0one run |
| 2× NVIDIA RTX 3080 20GBQwen3.8 27B at 4-bit | 20.0one run | 48.42 runs |
| 2× NVIDIA RTX 5060 Ti 16GBQwen3.8 27B at 3-bit | 29.2one run | 40.0one run |
Every multi-card setup
Plain runs on same-card setups, the most reported first. Reports that batched requests, used speculative decoding, mixed cards or did not say are left out here, and each is on its card's page.
2× NVIDIA RTX 309048 GB in all · 4 runs
- Qwen3.8 27B · 4-bit55.0 t/s (2 runs)
- Qwen3.8 27B · 8-bit37.1 t/s (2 runs)
2× AMD Radeon AI PRO R9700 32GB64 GB in all · 3 runs
- Qwen3.6 27B · 5-bit24.9 t/s (one run)
- Qwen3.8 125B · 6B active · 4-bit35.0 t/s (one run)
- Qwen3.8 27B · 4-bit37.0 t/s (one run)
4× AMD Radeon Pro V620 32GB128 GB in all · 3 runs
- Qwen3.8 125B · 6B active · 4-bit70.0 t/s (2 runs)
- Qwen3.8 27B · 8-bit30.0 t/s (one run)
2× NVIDIA Tesla P100 16GB32 GB in all · 3 runs
- Qwen3.8 27B · 6-bit36.3 t/s (2 runs)
- Qwen3.8 27B · 4-bit24.5 t/s (one run)
4× AMD Radeon AI PRO R9700 32GB128 GB in all · 2 runs
- Qwen3.8 125B · 6B active · 8-bit150.0 t/s (one run)
- Qwen3.8 27B · 8-bit36.6 t/s (one run)
2× NVIDIA RTX 3060 12GB24 GB in all · 2 runs
- Qwen3.8 27B · 4-bit36.4 t/s (2 runs)
2× NVIDIA RTX 3080 20GB40 GB in all · 2 runs
- Qwen3.8 27B · 4-bit48.4 t/s (2 runs)
4× NVIDIA RTX 309096 GB in all · 2 runs
- Qwen3.8 125B · 6B active · 4-bit55.0 t/s (2 runs)
2× NVIDIA RTX 5060 Ti 16GB32 GB in all · 2 runs
- Nemotron 3.5 Lightning 30B · 3B active · 4-bit80.0 t/s (one run)
- Qwen3.8 27B · 3-bit40.0 t/s (one run)
2× AMD MI50 16GB32 GB in all · one run
- Qwen3.6 35B · 3B active · 5-bit49.2 t/s (one run)
3× AMD MI50 16GB48 GB in all · one run
- Qwen3.8 27B · 8-bit30.1 t/s (one run)
2× AMD RX 7900 XT40 GB in all · one run
- Gemma 4 31B · 4-bit43.0 t/s (one run)
6× AMD RX 7900 XTX144 GB in all · one run
- Qwen3.8 27B · 16-bit50.0 t/s (one run)
2× AMD RX 9070 XT 16GB32 GB in all · one run
- Qwen3.8 27B · 4-bit24.7 t/s (one run)
2× AMD Strix Halo 128GB256 GB in all · one run
- Qwen3.8 125B · 6B active · 8-bit15.0 t/s (one run)
4× M3 Ultra 512GB2048 GB in all · one run
- Kimi K3 2800B · 104B active · 4-bit2.0 t/s (one run)
2× M4 16GB32 GB in all · one run
- Qwen3.6 35B · 3B active · 4-bit25.0 t/s (one run)
2× M5 Max 128GB256 GB in all · one run
- GLM-5.3 320B · 18B active · 1-bit24.0 t/s (one run)
2× NVIDIA A100 80GB160 GB in all · one run
- Qwen3.8 125B · 6B active · 2-bit56.8 t/s (one run)
8× NVIDIA A40 48GB384 GB in all · one run
- DeepSeek V4.1 Flash 552B · 16B active · 2-bit41.1 t/s (one run)
8× NVIDIA B300 288GB2304 GB in all · one run
- Kimi K3 2800B · 104B active · 4-bit92.0 t/s (one run)
4× NVIDIA CMP 170HX 64GB (unlocked)256 GB in all · one run
- DeepSeek V4 Flash 284B · 13B active · 4-bit29.0 t/s (one run)
2× NVIDIA DGX Spark256 GB in all · one run
- Nemotron-3-Ultra 550B · 55B active · 2-bit5.2 t/s (one run)
2× NVIDIA GTX 1080 Ti22 GB in all · one run
- Qwen3.6 35B · 3B active · 4-bit19.7 t/s (one run)
3× NVIDIA RTX 3060 12GB36 GB in all · one run
- Qwen3.8 27B · 6-bit25.0 t/s (one run)
3× NVIDIA RTX 309072 GB in all · one run
- Qwen3.8 125B · 6B active · 3-bit110.0 t/s (one run)
6× NVIDIA RTX 3090144 GB in all · one run
- Qwen3.8 125B · 6B active · 4-bit33.7 t/s (one run)
10× NVIDIA RTX 3090240 GB in all · one run
- DeepSeek V4 Flash 284B · 13B active · 8-bit7.2 t/s (one run)
4× NVIDIA RTX 5060 Ti 16GB64 GB in all · one run
- Qwen3.8 27B · 8-bit20.0 t/s (one run)
2× NVIDIA RTX 5070 Ti32 GB in all · one run
- Ornith1.5 35B · 3B active · 4-bit180.0 t/s (one run)
4× NVIDIA RTX 5070 Ti64 GB in all · one run
- Qwen3.8 27B · 8-bit82.2 t/s (one run)
2× NVIDIA RTX Pro 6000 Blackwell192 GB in all · one run
- DeepSeek V4 Flash 284B · 13B active · 8-bit106.1 t/s (one run)
4× NVIDIA RTX Pro 6000 Blackwell384 GB in all · one run
- GLM-5.3 320B · 18B active · 4-bit63.7 t/s (one run)
8× NVIDIA RTX Pro 6000 Blackwell768 GB in all · one run
- Tencent-HY3 295B · 21B active · 6-bit57.4 t/s (one run)
2× NVIDIA RTX PRO 6000 Max-Q192 GB in all · one run
- DeepSeek V4 Flash 284B · 13B active · 8-bit35.7 t/s (one run)
2× NVIDIA T4 16GB32 GB in all · one run
- Qwen2.5 7B · 4-bit58.0 t/s (one run)
3× NVIDIA Tesla P100 16GB48 GB in all · one run
- Qwen3.6 35B · 3B active · 4-bit70.0 t/s (one run)
2× NVIDIA V100 16GB32 GB in all · one run
- Qwen3.8 27B · 4-bit68.6 t/s (one run)
12× NVIDIA V100 32GB384 GB in all · one run
- Qwen3.6 35B · 3B active · 8-bit82.0 t/s (one run)
Frequently asked
Does a second GPU make a model run faster?
It depends on how the engine splits the model. Splitting by layers, llama.cpp's default, works the cards one after another, so it mostly buys memory for a bigger model. Tensor parallelism, in vLLM or SGLang, has every card work on each token at once and can raise the speed of a model that already fit, at the cost of traffic between the cards. The one-card and multi-card columns above show what happened on real setups.
What decides multi-GPU speed?
The engine and how it splits the model, the link between the cards (PCIe lanes, NVLink where a card has it), and whether every card is the same. These pages list same-card setups only, because a mixed pair runs at the pace of its slower card.