llamaperf

M3 Pro 36GB

APPLE · 36GB unified memory · 1 report

See what fits on this GPU →

Use the calculator to check which models fit in 36 GB of unified memory. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →All M3 Macs compared →
This page is thin (1 of 3 reports needed for indexing). Help fill it in.

Qwen3.8 27B

M3 Pro 36GB · custom C + Metal runtime

Tone: positive
reported speed:
17.7 tokens/s generation
quant:
Q4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingsummarization

User reports a custom C + Metal runtime with speculative decoding at 18 t/s end-to-end on coding tasks and 17.7 t/s for an LRUCache implementation, on 36 GB unified memory. Q4 weights are about 15 GB, mmap'd. Prose runs at roughly 10-11 t/s. TTFT is about 1.4s for a short prompt and about 2.7s for a 128-token prompt. The user compares the custom runtime against llama.cpp and reports it faster.

Sep 7, 2026