llamaperf

Running a local LLM on two or more GPUs

What people run on two or more cards of the same kind, the speeds they report, and the same card on its own beside it.

Can you use two GPUs for a local LLM?

Yes. People run local models on two cards and on rigs of up to 12, and the setup reported most is 2× NVIDIA RTX 3090. The usual reason is memory: two cards hold a model one cannot. Where the same card also ran the same model alone, the multi-card median was faster in all 4 cases, though that is few pairs and depends on the engine splitting the model across cards.

The calculator shows what fits on one card first, which tells you whether you need a second.

One card or several, same model

Median tokens per second for the same model size and quant level on the same card, alone and in a multi-card setup. Plain runs only: one request, no speculative decoding, nothing in system RAM.

Setup and modelOne cardSeveral
2× NVIDIA RTX 3090Qwen3.8 27B at 4-bit36.43 runs55.02 runs
2× AMD Radeon AI PRO R9700 32GBQwen3.8 27B at 4-bit29.0one run37.0one run
2× NVIDIA RTX 3080 20GBQwen3.8 27B at 4-bit20.0one run48.42 runs
2× NVIDIA RTX 5060 Ti 16GBQwen3.8 27B at 3-bit29.2one run40.0one run

Every multi-card setup

Plain runs on same-card setups, the most reported first. Reports that batched requests, used speculative decoding, mixed cards or did not say are left out here, and each is on its card's page.

2× NVIDIA RTX 309048 GB in all · 4 runs

2× AMD Radeon AI PRO R9700 32GB64 GB in all · 3 runs

4× AMD Radeon Pro V620 32GB128 GB in all · 3 runs

2× NVIDIA Tesla P100 16GB32 GB in all · 3 runs

4× AMD Radeon AI PRO R9700 32GB128 GB in all · 2 runs

2× NVIDIA RTX 3060 12GB24 GB in all · 2 runs

2× NVIDIA RTX 3080 20GB40 GB in all · 2 runs

4× NVIDIA RTX 309096 GB in all · 2 runs

2× AMD MI50 16GB32 GB in all · one run

3× AMD MI50 16GB48 GB in all · one run

2× AMD RX 7900 XT40 GB in all · one run

6× AMD RX 7900 XTX144 GB in all · one run

2× AMD RX 9070 XT 16GB32 GB in all · one run

2× AMD Strix Halo 128GB256 GB in all · one run

4× M3 Ultra 512GB2048 GB in all · one run

2× M4 16GB32 GB in all · one run

2× M5 Max 128GB256 GB in all · one run

2× NVIDIA A100 80GB160 GB in all · one run

8× NVIDIA B300 288GB2304 GB in all · one run

2× NVIDIA GTX 1080 Ti22 GB in all · one run

3× NVIDIA RTX 3060 12GB36 GB in all · one run

3× NVIDIA RTX 309072 GB in all · one run

6× NVIDIA RTX 3090144 GB in all · one run

4× NVIDIA RTX 5060 Ti 16GB64 GB in all · one run

2× NVIDIA RTX 5070 Ti32 GB in all · one run

4× NVIDIA RTX 5070 Ti64 GB in all · one run

2× NVIDIA T4 16GB32 GB in all · one run

3× NVIDIA Tesla P100 16GB48 GB in all · one run

2× NVIDIA V100 16GB32 GB in all · one run

12× NVIDIA V100 32GB384 GB in all · one run

Frequently asked

Does a second GPU make a model run faster?

It depends on how the engine splits the model. Splitting by layers, llama.cpp's default, works the cards one after another, so it mostly buys memory for a bigger model. Tensor parallelism, in vLLM or SGLang, has every card work on each token at once and can raise the speed of a model that already fit, at the cost of traffic between the cards. The one-card and multi-card columns above show what happened on real setups.

What decides multi-GPU speed?

The engine and how it splits the model, the link between the cards (PCIe lanes, NVLink where a card has it), and whether every card is the same. These pages list same-card setups only, because a mixed pair runs at the pace of its slower card.