Qwen3.8 27B
M5 Max 36GB · Inco Splash
- reported speed:
- 144.0 tokens/s generation
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3.8 27B at 144 t/s on an M5 Max MacBook Pro with 36 GB. Setup is the Inco Splash inference engine, an open-source engine built for Apple silicon, also available as a runtime inside LM Studio. Requirements are M3 or newer, macOS 26.4+, and 36 GB. The engine is claimed to reach up to 3x the decode speed of Ollama, 2x oMLX, and almost 4x when an agent fans out into sub-agents.