Gemma 4 26B (4B active)
2× NVIDIA GTX 1080 Ti · llama.cpp
- reported speed:
- 52.0 tokens/s generation · 340.9 tokens/s prompt processing
- quant:
- Q6_K_XL (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Gemma 4 26B-A4B at 52.03 t/s generation and 340.86 t/s prompt processing on a dual-GPU setup of GTX 1080 Ti and Radeon MI50 16GB, totaling 27GB VRAM. Setup is llama.cpp Vulkan pre-built binary build 851cb34f2 (11055) with the UD-Q6_K_XL GGUF, flash attention on, and 99 GPU layers offloaded. Single-GPU GTX 1080 Ti runs of the same model reached 11.76 t/s generation and 162.12 t/s prompt processing. The user also benchmarked Qwen3.6-35B-A3B MXFP4 MoE, Nemotron 31B-A3.5B Q5_K_M, Qwen3.8 27B Q6_K, and medgemma 27B Q6_K_XL, noting dense models benefited most from the second GPU.