Gemma 4 26B (4B active)
M6 32GB · SwiftLM
- reported speed:
- 52.2 tokens/s generation
- quant:
- 4-bit (MLX)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Gemma-4-26B-A4B 4-bit at 52.2 tok/s decode on a base Mac mini M6 with 32 GB unified memory. Setup is SwiftLM with MLX 4-bit weights, full GPU offload, peak GPU memory 19.5 GB; the longest prompt that passed was 80.7K tokens at 622 tok/s prefill and 24.3 tok/s decode. Also benchmarks Qwen3.6-35B-A3B 4-bit at 46.7 tok/s (GPU) and 13.2 tok/s (--stream-experts), Qwen3.8-27B 4-bit dense at 9.3 tok/s with 200 tok/s prefill, and Gemma-4-26B-A4B 8-bit at 8.8 tok/s with --stream-experts. An A/B against the prior revision shows 998 tok/s prefill at ~2.3K tokens versus 508 tok/s with swap, and 914 tok/s at ~9.5K tokens where the earlier build aborted on swap.