llamaperf

M4 Max 128GB

APPLE · 128GB unified memory · 5 reports

See what fits on this GPU →

Use the calculator to check which models fit in 128 GB of unified memory. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →All M4 Macs compared →
reported speed:
25.0 tokens/s generation
quant:
UD-Q2_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports running a model in-browser with custom WebGPU kernels, at speed comparable to llama.cpp.

Sep 7, 2026
reported speed:
72.5 tokens/s generation · 1472.4 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks Qwen3.6-35B-A3B on a Mac with OMLX, comparing MTP enabled against disabled. MTP shows minimal speedup for the 35B MoE model but roughly 2x for the 27B dense model. Results include pp and tg t/s at various context lengths and batch sizes.

Sep 7, 2026
reported speed:
15.2 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

long-contextagentic

User benchmarks Qwen3.8 27B on an M4 Max 128GB across five runtime and quant stacks at contexts from 32K to 256K. The stacks are oMLX AWQ 5-bit, oMLX oQ8e, mlx-dspark, and MTPLX 4-bit and 8-bit. Best decode at 256K is oMLX AWQ 5-bit at 15.2 t/s, while MTPLX collapses at 256K. Prefix caching is crucial for effective speed.

Sep 7, 2026
Tone: positive

User reports GLM 5.3 Flash support in the ds4 branch, running on an M4 Max with 128 GB. No performance numbers are provided.

Aug 28, 2026
Tone: positive
quant:
mixed-4_8bit (mlx)
codingmath

User reports Qwen 3.8 27B as the first model to break 94% on the cupel benchmark. Setup is llama.cpp with the unsloth/Qwen3.8-Flash-Next-GGUF:UD-IQ4_XS quant, which outperformed other 4-bit quants. Qwen 3.8 27B outperformed in coding but lost in general knowledge to Gemma 31B and Qwen 3.6.

Aug 28, 2026