Qwen3.8 27B
on M4 Max 128GB · oMLX · 262,144 ctx
Sep 29, 2026
Summary
User reports Qwen3.8 27B at 53.3 tok/s on a Mac Studio M4 Max 128GB.
Setup is oMLX 0.6.3rc2 with oQ4e 4-bit affine weights, 262,144 context, ANE prefill and native MTP speculative decoding with k=3 draft tokens, single stream.
Code generation measured 72.1 tok/s and prefill at 4K 273.7 tok/s. oMLX 0.6.1 with MTP gave 48.0 tok/s prose and 65.5 tok/s code; oMLX 0.6.3rc2 with ANE off gave 47.9 tok/s for both; vllm-metal 0.3.0 at 8-bit without MTP gave 13.2 tok/s code.