llamaperf

M5 Max 36GB

APPLE · 36GB unified memory · 1 report

See what fits on this GPU →

Use the calculator to check which models fit in 36 GB of unified memory. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →All M5 Macs compared →
This page is thin (1 of 3 reports needed for indexing). Help fill it in.

Qwen3.8 27B

M5 Max 36GB · Inco Splash

Tone: positive
reported speed:
144.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticcoding

User reports Qwen3.8 27B at 144 t/s on an M5 Max MacBook Pro with 36 GB. Setup is the Inco Splash inference engine, an open-source engine built for Apple silicon, also available as a runtime inside LM Studio. Requirements are M3 or newer, macOS 26.4+, and 36 GB. The engine is claimed to reach up to 3x the decode speed of Ollama, 2x oMLX, and almost 4x when an agent fans out into sub-agents.

Sep 19, 2026