llamaperf

M2 Max 64GB

APPLE · 64GB unified memory · 1 report

See what fits on this GPU →

Use the calculator to check which models fit in 64 GB of unified memory. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →All M2 Macs compared →
This page is thin (1 of 3 reports needed for indexing). Help fill it in.
reported speed:
2.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports DeepSeek V4.1 Flash at 1.8-2.2 t/s on an M2 Max 64GB. Setup is the Whallm app with MLX, loading parts of the model from SSD, around 33 GiB peak MLX memory, with 1K-16K input tokens. The app also adds built-in throughput benchmarks and an OpenAI-compatible API.

Sep 14, 2026