Gemma 4 26B (4B active)
on 2× NVIDIA GTX 1080 Ti · llama.cpp
Sep 19, 2026
Summary
User reports Gemma 4 26B-A4B at 52.03 t/s generation and 340.86 t/s prompt processing on a dual-GPU setup of GTX 1080 Ti and Radeon MI50 16GB, totaling 27GB VRAM.
Setup is llama.cpp Vulkan pre-built binary build 851cb34f2 (11055) with the UD-Q6_K_XL GGUF, flash attention on, and 99 GPU layers offloaded.
Single-GPU GTX 1080 Ti runs of the same model reached 11.76 t/s generation and 162.12 t/s prompt processing. The user also benchmarked Qwen3.6-35B-A3B MXFP4 MoE, Nemotron 31B-A3.5B Q5_K_M, Qwen3.8 27B Q6_K, and medgemma 27B Q6_K_XL, noting dense models benefited most from the second GPU.