llamaperf

Qwen3.8 27B

on M4 Max 128GB · oMLX · 262,144 ctx

Sep 29, 2026
Throughput
53.3 t/s gen · 273.7 t/s pp
Quant
oQ4e (MLX)
System RAM
128 GB

Summary

User reports Qwen3.8 27B at 53.3 tok/s on a Mac Studio M4 Max 128GB. Setup is oMLX 0.6.3rc2 with oQ4e 4-bit affine weights, 262,144 context, ANE prefill and native MTP speculative decoding with k=3 draft tokens, single stream. Code generation measured 72.1 tok/s and prefill at 4K 273.7 tok/s. oMLX 0.6.1 with MTP gave 48.0 tok/s prose and 65.5 tok/s code; oMLX 0.6.3rc2 with ANE off gave 47.9 tok/s for both; vllm-metal 0.3.0 at 8-bit without MTP gave 13.2 tok/s code.