Gemma 4 26B (4B active)
AMD Radeon 780M · llama.cpp
- reported speed:
- 25.0 tokens/s generation · 208.7 tokens/s prompt processing
- quant:
- Q4_K_M (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Gemma4 26B Q4_K_M at ~25.00 t/s generation and ~208.72 t/s prompt processing on an AMD Radeon 780M iGPU. Setup is llama.cpp with Vulkan backend, Q4_K_M GGUF, -ngl 99, on a MINISFORUM UM890 Pro mini PC running Ubuntu 24.04. A CLI smoke test gave ~23.4 t/s generation and ~37.3 t/s prompt; a no-reasoning run gave ~24.4 t/s generation and ~117.6 t/s prompt. Ollama on the same box was around 4.5 t/s generation, roughly a 5x-6x uplift. Ollama's installed Gemma4 blob could not be loaded directly in upstream llama.cpp due to a tensor count mismatch.