llamaperf

M5 Ultra 96GB

APPLE · 96GB unified memory · 3 reports

As of 7 Oct 2026, the models most run on the M5 Ultra 96GB, with the median of plain runs (one device, one request, no speculative decoding):

Engines people use on it: mlx-serve 2

Run models on your M5 Ultra 96GB? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the M5 Ultra 96GB

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 96 GB of unified memory.

Qwen3.8 27B Swift-1.5

M5 Ultra 96GB · mlx-serve · 25,000 ctx

Tone: positive
reported speed:
113.7 tokens/s generation · 3191.0 tokens/s prompt processing
quant:
4.7bpw (MLX)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 Swift-1.5 at 113.7 tok/s decode and 3191 tok/s prefill on an M5 Ultra 96GB Mac Studio. Setup is mlx-serve 26.10.1 with a 4.7bpw MLX quantization, 107GB download, text-only, tested to 179,200 tokens of context. Decode falls to 81.1 tok/s after a 95k prompt, where prefill is 2928 tok/s. User compares against a llama.cpp IQ3_XXS build of the same model, which reaches 62.7 tok/s decode after 4k and 1427 tok/s prefill at 25k. Top-1 agreement with Swift BF16 is 91.0% across 680 held-out positions, against 84.1% for the llama.cpp build.

Oct 6, 2026
Tone: positive
reported speed:
42.7 tokens/s generation · 828.2 tokens/s prompt processing
quant:
4-bit-8bit mix (MLX)
kv:
8-bit

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagenticlong-context

User reports Qwen3.8 Flash-Next on a base M5 Ultra 96GB (64-core GPU) at a median 828.2 t/s prefill and 42.7 t/s decode per stream, with a single-stream max of 3,328 t/s prefill and 152 t/s decode. Setup is a custom mlx-serve build with continuous batching at 4-way concurrency, a 4-bit/8-bit mixed MLX quant with MTP, 4 x 128k context (512k total) and 8-bit KV cache using 90GB of unified memory. Aggregate throughput was ~3,200 t/s prefill and ~170.8 t/s decode at 4-way concurrency. The run covered 112M tokens (109M prompt, 3M generated) with an 89% cache hit rate across ~1,700 sub-agent calls.

Sep 23, 2026
reported speed:
15.0 tokens/s generation
quant:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agentic

User compares an M5 Ultra 96GB against an M5 Max 128GB for running Qwen3.8-27B at Q8, and reports about 15 t/s on the Ultra. The user also discusses Qwen3.8-Flash-Next, a multimodal MoE with 176B total parameters and about 6B active, and estimates it will not fit in 96GB. MLX and llama.cpp are mentioned as possible engines, but the user does not confirm using either.

Sep 7, 2026

Get a weekly email of new M5 Ultra 96GB reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Qwen3.8 27B Swift-1.5
M5 Ultra 96GB
4.7bpw
mlx-serve
25,000113.7 tokens/s
Qwen3.8 125B (6B active) Flash-Next
M5 Ultra 96GB
4-bit-8bit mix
mlx-serve
131,07242.7 tokens/s
Qwen3.8 27B
M5 Ultra 96GB
Q8
Engine not reported
Not reported15.0 tokens/s