Qwen3.8 27B
on M5 Pro 64GB · oMLX · 16,384 ctx
Sep 27, 2026
Summary
User reports Qwen3.8 27B at 22.99 tokens/s three-scenario decode average on an Apple M5 Pro with 64 GB of unified memory.
Setup is oMLX 0.6.0 with 4-bit affine MLX weights and an external Qwen3.8-27B-MTP-4bit drafter (250.9 MB, block size 3), measured at 16384 tokens of context.
The average combines 26.28 short, 22.81 long, and 19.88 at 11k context. First uncached 11k prefill was 368.32 tokens/s with 30.18 s TTFT. MTP accepted 1,836 of 3,416 drafted tokens (53.7%).