llamaperf

M5 Pro 48GB

APPLE · 48GB unified memory · 2 reports

See what fits on this GPU →

Use the calculator to check which models fit in 48 GB of unified memory. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →All M5 Macs compared →
This page is thin (2 of 3 reports needed for indexing). Help fill it in.
Tone: mixed
reported speed:
44.0 tokens/s generation
quant:
Q4_K_M (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingtool-usemath

User compares DeepSeek V4 Flash 9B distill against its base model Qwen3.5 9B on a MacBook Pro M5 Pro with 48 GB RAM. Both models gave the same answers on 6/8 tasks, but the distill used fewer tokens (5480 vs 8975). Throughput was near identical (44 vs 41 tok/s). The distill was slower on open-ended tasks (1467 tokens vs 867 for a buffer overflow explanation). The base model failed arithmetic due to a token cap.

Sep 7, 2026
Tone: mixed
reported speed:
19.0 tokens/s generation · 325.0 tokens/s prompt processing
quant:
Q8_0 (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User compares MLX and GGUF for Qwen3.8 27B on an M5 Pro, reporting generation at about 19 t/s. Setup is llama.cpp with Metal, which the user says now matches MLX prefill at roughly 300-350 t/s. The user questions whether MLX is still needed and mentions MTP and oMLX.

Sep 7, 2026