llamaperf

M3 Ultra 512GB

APPLE · 512GB unified memory · 4 reports

See what fits on this GPU →

Use the calculator to check which models fit in 512 GB of unified memory. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →All M3 Macs compared →

DeepSeek V4.1 Flash

M3 Ultra 512GB · oMLX · 65,536 ctx

Tone: positive
reported speed:
19.7 tokens/s generation · 439.4 tokens/s prompt processing
quant:
oQ4e
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports DeepSeek V4.1 Flash on an M3 Ultra 512GB with oQ4e and Engram in RAM, reaching 439.4 t/s prefill and 19.7 t/s generation at 64K context with MTP off, and 435.6 t/s prefill and 39.7 t/s generation with MTP on. Setup is oMLX 0.7.0.dev2 with Python code prompts, temperature 1.0, 128 generated tokens, and no prefix cache. One measured run per configuration after warm-up. At 4K context, prefill was 458.0 t/s with MTP off and 452.2 t/s with MTP on, while generation was 20.2 t/s and 32.1 t/s respectively. At 16K, prefill was 459.1 t/s and 454.8 t/s, with generation at 20.0 t/s and 34.7 t/s. At 32K, prefill was 452.2 t/s and 447.7 t/s, with generation at 19.8 t/s and 31.5 t/s. The user also notes experimental MoE Expert SSD Offload support for DeepSeek V4.1, Qwen3.8-Flash-Next, Gemma 4 MoE, and OLMoE, and an M5 prefill speedup from 615.2 to 826.7 t/s at 32K on an M5 Max with Qwen3.8-27B using INT8-activation kernels.

Sep 12, 2026
Tone: positive
reported speed:
20.0 tokens/s generation · 533.0 tokens/s prompt processing
quant:
8bit (affine)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports an optimized DeepSeek V4 Flash 8-bit affine MLX model on oMLX, with prefill improving from ~300-321 to ~533 tok/s and decode from ~7.31 to ~20-22 tok/s. Real runs at 79K-119K context show 19.2-20.7 tok/s. User asks for community review on accuracy and next optimization directions.

Sep 7, 2026
Tone: negative
reported speed:
8.0 tokens/s generation
quant:
Q4_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports 8 t/s on a Mac Studio M3 Ultra 512GB with DeepSeek V4 Flash GGUF Q4_K_XL. The user also tried Q8. The user expected better performance.

Sep 7, 2026
Tone: positive
reported speed:
475.0 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports DeepSeek V4 Flash serving on an M3 Ultra 512GB with the ds4 engine, with cold prefill improved from 392 to 475 t/s at 64k context through kernel optimizations. Cache prewarming with max_tokens:0 yields roughly 10x speedup for chat turns, from 6-20s down to 1.6s.

Sep 7, 2026