llamaperf

M3 Ultra 256GB

APPLE · 256GB unified memory · 2 reports

See what fits on this GPU →

Use the calculator to check which models fit in 256 GB of unified memory. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →All M3 Macs compared →
This page is thin (2 of 3 reports needed for indexing). Help fill it in.
Tone: negative
reported speed:
17.3 tokens/s generation
quant:
4bit (MLX)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports a bug in oMLX 0.6.4 distributed clustering where the coordinator fails to release RAM and GPU after a crash. The model is mlx-community/MiniMax-M3-4bit at 236 GB, run across rank 0 on a Mac Studio M3 Ultra 256GB and rank 1 on a Mac Studio M2 Ultra 192GB. The first completion produced 17 tokens from a 7,693-token prompt at 17.3 tok/s. The crash was triggered by overlapping requests with pipeline_prefill_overlap and coalesced batching. After the crash, roughly 116 GB of wired memory has no owning process and the GPU stays pinned at 100%, requiring a reboot. The M2 Ultra worker released memory cleanly. The user also reports that MiniMax-M3 crashes the cluster after the first prompt.

Sep 10, 2026
Tone: positive
reported speed:
37.4 tokens/s generation · 550.0 tokens/s prompt processing
quant:
Q4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticcodingtool-use

User reports GLM-5.3-Flash on an M3 Ultra at 60 t/s for SQL and 38 t/s average with an agent harness. Setup uses a dflash drafter for speculative decoding, with prefill at 550 t/s at 62k context and generation at 37.4 t/s at 300k context. The user states accuracy is preserved.

Sep 9, 2026