llamaperf

Qwen3.8 27B

on M5 Pro 64GB · oMLX · 16,384 ctx

Sep 27, 2026
Throughput
23.0 t/s gen · 368.3 t/s pp
Quant
4bit (MLX)
System RAM
64 GB

Summary

User reports Qwen3.8 27B at 22.99 tokens/s three-scenario decode average on an Apple M5 Pro with 64 GB of unified memory. Setup is oMLX 0.6.0 with 4-bit affine MLX weights and an external Qwen3.8-27B-MTP-4bit drafter (250.9 MB, block size 3), measured at 16384 tokens of context. The average combines 26.28 short, 22.81 long, and 19.88 at 11k context. First uncached 11k prefill was 368.32 tokens/s with 30.18 s TTFT. MTP accepted 1,836 of 3,416 drafted tokens (53.7%).