llamaperf

M5 32GB

APPLE · 32GB unified memory · 2 reports

See what fits on this GPU →

Use the calculator to check which models fit in 32 GB of unified memory. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →All M5 Macs compared →
This page is thin (2 of 3 reports needed for indexing). Help fill it in.
Tone: positive
reported speed:
1.0 tokens/s generation · 50.0 tokens/s prompt processing
quant:
4bit

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports DeepSeek V4 Flash 0731 at about 50 t/s prefill and about 1 t/s decode on an M5 Air with 32 GB. Setup uses the streamed experts trick with a model of roughly 300B total parameters.

Sep 7, 2026

Gemma 4

M5 32GB · MLX · 130,173 ctx

reported speed:
3029.0 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports custom w8a8 kernels on an M5 MacBook Air, with prefill improving from 2193 t/s to 3029 t/s for 130k tokens. User also mentions llama.cpp for Macs.

Jul 23, 2026