Gemma 4 5.1B E2B Instruct
M1 8GB · Ollama
- reported speed:
- 17.5 tokens/s generation
- quant:
- Q4 (gguf)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
text-generation
User reports Gemma 4 at 15-20 t/s, a range the user calls usable for simple tasks.