llamaperf

M2 Pro 32GB

APPLE · 32GB unified memory · 1 report

See what fits on this GPU →

Use the calculator to check which models fit in 32 GB of unified memory. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →All M2 Macs compared →
This page is thin (1 of 3 reports needed for indexing). Help fill it in.

Qwen3.8 27B

M2 Pro 32GB · llama.cpp · 131,072 ctx

Tone: mixed
reported speed:
8.6 tokens/s generation · 21.9 tokens/s prompt processing
quant:
IQ4_XS (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingvision

User reports Qwen3.8 27B IQ4_XS with vision projector on an M2 MacBook Pro 32GB, at 21.9 t/s prompt and 8.6 t/s generation. Setup is llama.cpp built from source, with context set to 128K due to memory constraints. User notes the model is slow on Mac but that thinking quality is good, and mentions previous experience with Qwen3.5 35B-A3B.

Sep 7, 2026