Gemma 4 26B (4B active) IT
M2 8GB · Turbo-fieldfare
- throughput:
- 5.5 t/s gen
Custom Swift/Metal engine. Also tested on M5 MacBook Pro: 31-35 tok/s. OpenAI-compatible server with streaming and tool-call support.
Google DeepMind · 22 reports
| Engine | Avg t/s | Range | N |
|---|---|---|---|
| Ollama | 61.2 | 18–150 | 6 |
| llama.cpp | 40.5 | 8–97 | 3 |
| vLLM | 318.0 | 58–578 | 2 |
M2 8GB · Turbo-fieldfare
Custom Swift/Metal engine. Also tested on M5 MacBook Pro: 31-35 tok/s. OpenAI-compatible server with streaming and tool-call support.
User built custom w8a8 kernels for M5 MacBook Air. Baseline prefill 2193 tps, improved to 3029 tps for 130k tokens. Also mentions llama.cpp for Macs.
RX 7900 XTX · llama-swap
Benchmark of Gemma 4 QAT vs regular quants on AMD 7900 XTX. No token/s reported, but wall clock times show significant speedups (e.g., 12B QAT 45% faster, 83% throughput increase). Quality reported identical. Models tested: 12B, 26B, 31B, E4B.
Benchmark of Gemma 4 26B-A4B vs 12B on RTX 4090. 26B-A4B used 15GB VRAM, 138 tok/s; 12B used 9GB, 80 tok/s. 26B-A4B won every scene and ran ~1.7x faster. 12B ideal for 16GB laptop.
RTX 5090 · llama.cpp
Benchmark of FWHT CUDA implementation for kv-cache quantization. Results show 1-2% pp boost and 7-9% tg boost on Gemma 4 26B.A4B Q4_K_M with -ctk q8_0 -ctv q8_0. pp2048 and tg128 values reported; highest t/s from cuda-fwt column.
H100 80GB · vLLM · 32,768 ctx
Benchmark of Gemma 4 31B dense with MTP and DFlash speculative decoding. Also tested Gemma 4 26B-A4B MoE (25.2B total, 3.8B active). MTP 3.11x faster, DFlash 3.03x faster than baseline at concurrency 1. Baseline 40.3 tok/s, MTP 125.3 tok/s, DFlash 122.1 tok/s. At concurrency 16: baseline 375 tok/s, MTP 953 tok/s, DFlash 725 tok/s. For MoE: baseline 177.1 tok/s, MTP 264.2 tok/s, DFlash 306.4 tok/s at concurrency 1. At concurrency 16: baseline 975 tok/s, MTP 1808 tok/s, DFlash 1957 tok/s. Coding, math, STEM, reasoning benefited more.
DFlash speculative decoding with vLLM 0.19.2rc1. Baseline 228 t/s, best with DFlash 578 t/s (2.56x speedup). Draft model: z-lab/gemma-4-26B-A4B-it-DFlash. Input 256 tokens, output 1024 tokens.
M5 Max 64GB · llama.cpp
Multi-Token Prediction (MTP) implementation yields 40% speedup (138 t/s with MTP).
User reports poor performance with Gemma 4 (7.5 tok/s) and Qwen3.6-27B (locking up), while Qwen3.6-35B-A3 is fast. Suspects a bug with dense models.
User reports poor performance with dense models (Gemma4-31B ~7.5 t/s, Qwen3.6-27B locking up) on M5 Max 128GB, while Qwen3.6-35B-A3B MoE is fast. Mentions using DFLASH (likely flash attention).
RX 7900 XTX · Ollama
~35-45 tok/s on RX 7900 XTX. Gemma 4 E4B Q4_K_M via Ollama. Best consumer AMD option. Source: gemma4-ai.com AMD GPU guide
Gemma 4 E4B on RX 7900 XTX via vLLM + ROCm. Default path: 57.96 gen tok/s, 82.96 prompt tok/s. Source: flexinfer.ai
User reports poor performance with Gemma4-31B (7.5 tok/s) and Qwen3.6-27B (locking up) on M5 Max 128GB, while Qwen3.6-35B-A3 is fast. Mentions using DFLASH.
AMD Threadripper 256GB · llama.cpp
~5-10 tok/s on CPU. E2B is usable CPU-only. Source: gemma4-ai.com hardware guide
RTX 4070 · Ollama
~55 tok/s on RTX 4070 12GB. Ada Lovelace efficiency. Source: estimated from compute-market tiers
~60 tok/s on RTX 3060 12GB. E2B runs effortlessly. Source: estimated from compute-market tiers
RTX 3060 12GB · Ollama
~45 tok/s on RTX 3060 12GB. E4B fits easily. Source: compute-market.com
RTX 5090 · 50,000 ctx
nvidia/Gemma-4-26B-A4B-NVFP4 works on 5090 with 80% allocation (of 32GB) got around 50k context. Model size 18.8GB. Benchmarks provided: GPQA Diamond 80.30% (baseline) vs 79.90% (NVFP4), AIME 2025 88.95% vs 90.00%, MMLU Pro 85.00% vs 84.80%, LiveCodeBench (pass@1) 80.50% vs 79.80%, IFBench 77.77% vs 78.1%, IFEval 96.60% vs 96.40%.
~150 tok/s generation. Star performer. Source: n1n.ai
RTX A6000 48GB · llama.cpp
bf16 no quantization. 10.25GB VRAM, 61ms TTFT. Source: dev.to Gaurav Vij
M1 8GB · Ollama
15-20 tok/s range, usable for simple tasks. Source: gemma4-ai.com
m5 pro (18 core cpu, 20 core gpu) · Ollama
User reports very fast performance on M5 Pro with 48GB RAM. Prompt eval rate: 204.07 t/s, generation rate: 68.76 t/s.