llamaperf

M4 Max 64GB

APPLE · 64GB unified memory · 2 reports

See what fits on this GPU →

Use the calculator to check which models fit in 64 GB of unified memory. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →All M4 Macs compared →
This page is thin (2 of 3 reports needed for indexing). Help fill it in.
Tone: positive
reported speed:
11.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports DeepSeek-V4-Flash (284B) at about 11 t/s on a 64GB MacBook. The model is 165GB on disk and does not fit in RAM, so it runs via SSD streaming with a 2-bit file. Quality matches official perplexity (6.1250 vs 6.1262). Higher quality settings give 1-2 t/s. An 8GB cache gave 2.04 t/s versus 1.23 t/s with 32GB. Speculative decoding and prefetching did not help.

Sep 7, 2026
Tone: mixed
reported speed:
36.0 tokens/s generation · 517.9 tokens/s prompt processing
quant:
AD-3.84bpw-M64 (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports a 85 GB GGUF quant of Qwen 3.8 Flash Next running on a 64 GB MacBook with the ngram table offloaded to SSD, at 517.9 t/s prefill and 36 t/s decode. The user notes the quant quality is far from perfect and better versions are planned.

Sep 7, 2026