Qwen3.8 27B
on M4 Pro 48GB · oMLX · 200,000 ctx
Sep 29, 2026
Summary
User reports Qwen3.8-27B at 24.8 tok/s decode on an M4 Pro 48GB, with median TTFT 943 ms over three samples of 4096 prompt tokens and 128 generated tokens.
Setup is oMLX 0.6.4 with the oQ4e 4-bit MLX model and Lightning MTP draft depth 3 enabled via the mtp3 profile, at a 200k context window.
Enabling MTP through a profile rather than settings.json raised decode from 10 tok/s to 24.8 tok/s. A separate sweep documents the MTP-on prefill curve at 1.8k-30k prompts and a 5-turn agentic session reaching 157k context with 85% token reuse.