llamaperf

M5 32GB

APPLE · 32GB unified memory · 1 report

See what fits on this GPU →
This page is thin (1 of 3 reports needed for indexing). Help fill it in.

Gemma 4

M5 32GB · MLX · 130,173 ctx

throughput:
3029.0 t/s pp

User built custom w8a8 kernels for M5 MacBook Air. Baseline prefill 2193 tps, improved to 3029 tps for 130k tokens. Also mentions llama.cpp for Macs.