llamaperf

M1 Max 64GB

APPLE · 64GB unified memory · 4 reports

See what fits on this GPU →

Use the calculator to check which models fit in 64 GB of unified memory. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →All M1 Macs compared →
Tone: positive
reported speed:
8.0 tokens/s generation · 30.0 tokens/s prompt processing
quant:
IQ3-XXS

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports DeepSeek-V4-Flash-0731 at about 8 t/s decode and about 30 t/s prefill on an M1 Max 64GB. Setup is a patched llama.cpp with the IQ3-XXS quant at 104 GB and context limited to 64k.

Sep 7, 2026
reported speed:
10.0 tokens/s generation · 165.0 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports a run on a 64GB M1 Max. No drafter was used. User mentions n-gram and stripping embeddings as potential optimizations.

Sep 7, 2026

Qwen3.8 27B

M1 Max 64GB · MTPLX · 262,000 ctx

Tone: positive
reported speed:
21.0 tokens/s generation · 83.0 tokens/s prompt processing
quant:
Q4 (mlx)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports a roughly 2x speed boost for Qwen3.8 27B on Apple Silicon using the MTPLX framework. User also reports Qwen3.6 35B A3B at about 55 t/s decode and about 300 t/s prefill, with a peak of 623 t/s.

Sep 7, 2026
Tone: positive
reported speed:
180.0 tokens/s prompt processing
quant:
Q4 (gguf)
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports a custom Q4 quant with spliced tensors from Unsloth and AtomicChat quants, with MTP enabled giving +70% decode at 22 btps. Setup uses custom metal-optimized sparse attention, with prefill reduced to 170 tps at 4K and 150 tps at 256K context. Q4_0 MTP matches Q8_0 acceptance rates at half RAM.

Sep 7, 2026