llamaperf

M1 Max 32GB

APPLE · 32GB unified memory · 2 reports

See what fits on this GPU →

Use the calculator to check which models fit in 32 GB of unified memory. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →All M1 Macs compared →
This page is thin (2 of 3 reports needed for indexing). Help fill it in.
reported speed:
15.8 tokens/s generation · 81.8 tokens/s prompt processing
quant:
4bit (MLX)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User compares MLX and llama.cpp on Qwen3.8-27B. MLX uses mlx-community/Qwen3.8-27B-4bit at ~16.1GB, with prompt processing at 81.76 tok/s, generation at 15.81 tok/s, and peak memory of 16.39GB. llama.cpp uses unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M at 15.32 GiB and 27.32B params, with full Metal offload and Flash Attention enabled, giving prompt processing at 99.61 ± 0.44 tok/s and generation at 9.69 ± 0.34 tok/s. llama.cpp is ~22% faster at prompt processing, while MLX is ~63% faster at generation.

Sep 10, 2026
reported speed:
15.8 tokens/s generation · 81.8 tokens/s prompt processing
quant:
4bit

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User compares MLX and llama.cpp on Qwen3.8-27B. MLX reaches 15.81 t/s generation and 81.76 t/s prompt processing. llama.cpp reaches 9.69 t/s generation and 99.61 t/s prompt processing. The llama.cpp run uses the UD-Q4_K_M quant.

Sep 9, 2026