llamaperf

M5 Max 64GB

APPLE · 64GB unified memory · 4 reports

See what fits on this GPU →

Use the calculator to check which models fit in 64 GB of unified memory. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →All M5 Macs compared →
Tone: positive
reported speed:
30.0 tokens/s generation
quant:
Q2-Q4 mixed imatrix (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticcoding

User reports about 30 t/s on an M5 Max with a quantized DeepSeek V4 Flash running on the DS4 engine. The user states this is double the speed of llama.cpp. The user asks for feedback from CUDA and ROCm users.

Sep 7, 2026
Tone: positive
reported speed:
97.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports a Multi-Token Prediction (MTP) implementation yielding a 40% speedup, reaching 138 t/s with MTP.

May 8, 2026
Tone: positive
reported speed:
63.0 tokens/s generation
quant:
4-bit (mlx)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingcreative-writing

User reports 63 t/s on Qwen3.6-27B 4-bit MLX on an M5 Max 64GB with the MTPLX engine, up from a 28 t/s baseline. Setup is a custom patched MLX fork with Metal kernels, using native MTP heads at temperature 0.6, top_p 0.95 and top_k 20, with optimal depth D3.

May 5, 2026
Tone: mixed
reported speed:
32.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen 3.6 27B at 32 t/s on a MacBook Pro M5 Max 64GB, generating 33,946 tokens in 18m04s. User compares Gemma 4 31B on the same hardware at 27 t/s, generating 6,209 tokens in 3m51s. User finds Qwen showed more creativity while Gemma won for game logic.

May 1, 2026