llamaperf

M5 Pro 24GB

APPLE · 24GB unified memory · 3 reports

As of 7 Oct 2026, the models most run on the M5 Pro 24GB, with the median of plain runs (one device, one request, no speculative decoding):

Engines people use on it: MLX 1 · llama.cpp 1

Run models on your M5 Pro 24GB? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the M5 Pro 24GB

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 24 GB of unified memory.

Tone: positive
reported speed:
73.9 tokens/s generation
quant:
IQ3_XXS (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports decode speed improvements on Qwen3-30B-A3B with IQ3_XXS quantization on an M5 MacBook Pro 24GB. The decode speed increased from 65.6 to 73.9 tok/s using a llama.cpp PR for Metal optimization. Setup uses llama.cpp with IQ3_XXS quantization. The user also tested on an M1 Pro 32GB. The user is looking for testers for Qwen3.8-Flash-Next and provides benchmark commands.

Oct 6, 2026
Tone: positive
reported speed:
125.8 tokens/s generation · 363.9 tokens/s prompt processing
quant:
4bit (MLX)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagentic

User benchmarks DeepSeek-Coder-V2-Lite-Instruct at 125.8 tok/s generation and 363.9 tok/s prompt processing on an Apple M5 Pro with 24GB unified memory. Setup is MLX with a 4bit quant (8.2GB) on a single machine, running a fixed coding task capped at 1500 output tokens. The model produced correct working code; the user notes that 24GB has a hard ceiling around 20GB models and that background downloads measurably affect tok/s.

Oct 3, 2026
Tone: positive
reported speed:
4.8 tokens/s generation
quant:
2-bit dynamic

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports DeepSeek-V4-Flash 284B-A13B at 4.8 tok/s on a 24 GB M5 Pro with about 5.3 GB of memory in use. Setup is Mference, an open-source engine. The user also runs Gemma 4 26B-A4B and Qwen 3.6 35B-A3B.

Sep 7, 2026

Get a weekly email of new M5 Pro 24GB reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Qwen3 30B (3B active)
M5 Pro 24GB
IQ3_XXS
llama.cpp
Not reported73.9 tokens/s
DeepSeek-Coder-V2-Lite 16B (2.4B active)
M5 Pro 24GB
4bit
MLX
Not reported125.8 tokens/s
DeepSeek V4 Flash 284B (13B active)
M5 Pro 24GB
2-bit dynamic
Mference
Not reported4.8 tokens/s