llamaperf

M3 Max 96GB

APPLE · 96GB unified memory · 2 reports

See what fits on this GPU →

Use the calculator to check which models fit in 96 GB of unified memory. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →All M3 Macs compared →
This page is thin (2 of 3 reports needed for indexing). Help fill it in.
Tone: negative
reported speed:
7.4 tokens/s generation · 121.3 tokens/s prompt processing
quant:
8bit (MLX)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.6 27B at 7.4 t/s generation and 121.3 t/s prompt processing on an M3 Max 96GB. Setup is oMLX with the MLX 8-bit model at pp1024/tg128, using 28.34 GB peak memory. A pp4096/tg128 run reached 8.8 t/s generation and 133.8 t/s prompt processing. Continuous batching at 4x reached 19.9 t/s aggregate generation. User is new to LLMs and asks whether the slow speed is due to the dense model or a setup problem.

Sep 12, 2026
Tone: positive
reported speed:
12.7 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports DeepSeek V4 Flash at 11-13 t/s on an M3 Max 96GB with SSD streaming and iogpu.wired_limit_mb=86016. Setup is antirez's ds4 engine with GGUF. TTFT is 3-5s after warmup. Prefill of 36k tokens takes about 2.5 minutes.

Sep 7, 2026