llamaperf

AMD Radeon 780M

AMD · shared memory · 5 reports

As of 7 Oct 2026, the models most run on the AMD Radeon 780M, with the median of plain runs (one device, one request, no speculative decoding):

Engines people use on it: llama.cpp 4 · Ollama 1

Run models on your AMD Radeon 780M? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the AMD Radeon 780M

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. It has no memory of its own and borrows the machine's RAM, so how large a model fits depends on the machine, and the calculator does not size it.

Tone: positive
reported speed:
25.0 tokens/s generation · 208.7 tokens/s prompt processing
quant:
Q4_K_M (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Gemma4 26B Q4_K_M at ~25.00 t/s generation and ~208.72 t/s prompt processing on an AMD Radeon 780M iGPU. Setup is llama.cpp with Vulkan backend, Q4_K_M GGUF, -ngl 99, on a MINISFORUM UM890 Pro mini PC running Ubuntu 24.04. A CLI smoke test gave ~23.4 t/s generation and ~37.3 t/s prompt; a no-reasoning run gave ~24.4 t/s generation and ~117.6 t/s prompt. Ollama on the same box was around 4.5 t/s generation, roughly a 5x-6x uplift. Ollama's installed Gemma4 blob could not be loaded directly in upstream llama.cpp due to a tensor count mismatch.

Sep 26, 2026
Tone: mixed
reported speed:
3.2 tokens/s generation · 32.5 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen3.8 27B at 3.20 t/s generation and 32.45 t/s prompt processing on a Ryzen 9 7940HS with Radeon 780M iGPU and 24GB RAM. Setup is Ollama on two machines: a Minisforum UM790 Pro and a home build with the same Ryzen 9 7940HS, both with 24GB RAM. The Minisforum run reached 4.73 t/s generation and 20.52 t/s prompt processing. Both runs stopped mid-task and follow-up prompts returned no output. User expected low throughput but not the failures.

Sep 18, 2026
Tone: positive
reported speed:
21.1 tokens/s generation · 287.3 tokens/s prompt processing
quant:
Q8_0
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks Qwen3.6 35B-A3B Q8_0 on a Ryzen 7 260 with a 780M iGPU and 64 GB DDR5, reaching pp8192 287.33 t/s and tg128 21.06 t/s. Setup is the Vulkan backend. Gemma 4 31B Q8_0 on the same hardware gives pp8192 51.59 t/s and tg128 2.46 t/s. With MTP, Gemma 4 31B reaches tg ~5.76 t/s, and Qwen3.6 35B-A3B with MTP and partial offloading reaches tg ~34.85 t/s. The user mentions an RTX 5060 8GB as a bonus for MoE partial offloading.

Sep 7, 2026
reported speed:
18.4 tokens/s generation · 311.4 tokens/s prompt processing
quant:
Q8 (GGUF)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks ROCm against Vulkan on a Radeon 780m iGPU. ROCm gives a 50% pp speedup for dense models. The user also reports Qwen3.8 27B results.

Sep 7, 2026
Tone: mixed
reported speed:
100.0 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports ROCm 7.14 on a Radeon 780M with llama.cpp is fast but unstable and crashes. Setting AMD_SERIALIZE_KERNEL=3 stabilizes it but drops prompt processing to about 100 t/s, down from 200-300 t/s without the workaround. User asks for advice.

Sep 7, 2026

Get a weekly email of new AMD Radeon 780M reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Gemma 4 26B (4B active)
AMD Radeon 780M
Q4_K_M
llama.cpp
Not reported25.0 tokens/s
Qwen3.8 27B
2× AMD Radeon 780M
Not reported
Ollama
Not reported3.2 tokens/s
Qwen3.6 35B (3B active)
AMD Radeon 780M
Q8_0
llama.cpp
Not reported21.1 tokens/s
Qwen3.6 35B (3B active)
AMD Radeon 780M
Q8
llama.cpp
Not reported18.4 tokens/s