Gemma 4 26B (4B active)
NVIDIA RTX 5060 Ti 16GB · llama.cpp · 65,536 ctx
- reported speed:
- 95.2 tokens/s generation · 3476.0 tokens/s prompt processing
- quant:
- MXFP4-MOE (GGUF)
- kv:
- q4_0
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Gemma 4 26B-A4B-it at 95.22 tok/s decode and 3,476 tok/s prompt processing on an RTX 5060 Ti 16 GB. Setup is a custom llama.cpp build (commit 0b484ab2b plus 12 commits) with MXFP4-MOE quantization (experts MXFP4, dense Q8_0, 4.66 BPW) and q4_0 KV cache, 1 slot, flash attention enabled, 65,536 token context. The MXFP4-MOE quant brings the model to 13.70 GiB, fitting the 16 GB card where Q4_K_M at 15.85 GiB does not. On an RTX 5090 the same quant gives 10,733 tok/s prompt and 196.6 tok/s decode versus 8,744 and 219.9 for Q4_K_M, with perplexity 3,864 vs 3,615.