llamaperf

Gemma 4 8B E4B

on M1 16GB · Unsloth Studio · 131,100 ctx

Tone: positive
Oct 4, 2026
Throughput
8.2 t/s gen
Quant
4bit (MLX)
System RAM
16 GB

Use cases

codingagenticsummarization

Summary

User reports Gemma 4 E4B at around 8.2 tok/s on an M1 MacBook Air with 16GB unified memory. Setup is Unsloth Studio with a 4-bit MLX model and a 131.1k context window. The user says this speed is fine for chatbot use and asks for suggestions of 7B or 8B MLX models for coding and background tasks.