llamaperf

M5 Pro 64GB

APPLE · 64GB unified memory · 4 reports

See what fits on this GPU →

Use the calculator to check which models fit in 64 GB of unified memory. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →All M5 Macs compared →
Tone: mixed
reported speed:
20.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 27B at around 20 t/s sustained on an M5 Pro 64GB, with about 30 t/s in the first 2-4k tokens. User finds it unbearably slow, possibly because the model thinks so much. User also tried antirez's DS4 with DeepSeek V4 Flash and gets around 8 t/s sustained, which they call slow and basically unusable.

Sep 13, 2026
Tone: positive
reported speed:
15.0 tokens/s generation
quant:
Q8_0 (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports 10-17 t/s in SSD streaming mode. Setup keeps routed experts partly in RAM cache and pulls them from the GGUF on cache misses. Experts and output head are Q8_0, while router, embeddings and V4 auxiliary blocks are FP16.

Sep 7, 2026
reported speed:
10.5 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagentic

User asks which model to use for agentic coding. User reports DeepSeek V4 Flash at 10-11 t/s with 1M context, and Qwen3.8 27B at 15-16 t/s with FP8 quants.

Sep 7, 2026
Tone: mixed
reported speed:
2.0 tokens/s generation
quant:
mixed 4/8-bit

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports 2.04 t/s with an 8 GB expert cache and 1.23 t/s with a 32 GB expert cache, streaming from SSD. Speculative decoding and prefetch both decreased throughput. Codebook 2-bit quantization failed the quality gate.

Sep 7, 2026