llamaperf

Gemma 4 26B (4B active)

on AMD Radeon 780M · llama.cpp

Tone: positive
Sep 26, 2026
Throughput
25.0 t/s gen · 208.7 t/s pp
Quant
Q4_K_M (GGUF)

Summary

User reports Gemma4 26B Q4_K_M at ~25.00 t/s generation and ~208.72 t/s prompt processing on an AMD Radeon 780M iGPU. Setup is llama.cpp with Vulkan backend, Q4_K_M GGUF, -ngl 99, on a MINISFORUM UM890 Pro mini PC running Ubuntu 24.04. A CLI smoke test gave ~23.4 t/s generation and ~37.3 t/s prompt; a no-reasoning run gave ~24.4 t/s generation and ~117.6 t/s prompt. Ollama on the same box was around 4.5 t/s generation, roughly a 5x-6x uplift. Ollama's installed Gemma4 blob could not be loaded directly in upstream llama.cpp due to a tensor count mismatch.