Gemma 4 26B (4B active)
M2 8GB · Turbo-fieldfare
- reported speed:
- 5.5 tokens/s generation
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports 31-35 t/s on an M5 MacBook Pro with a custom Swift/Metal engine. Setup is an OpenAI-compatible server with streaming and tool-call support.