Qwen3.8 27B
on M5 Max 128GB · mlx-dspark
Oct 7, 2026
Use cases
codingagentic
Summary
User reports Qwen 3.8 27B at 39.9 tok/s with DFlash 2 speculative decoding on an M5 Max 128GB, versus 17.9 tok/s plain decode.
Setup is mlx-dspark 0.14 on MLX 0.32 with the 8-bit MLX build (~29 GB) and 4-bit/8-bit quantized KV cache; the DFlash 2 drafter is incoai/Qwen3.8-27B-DFlash2 (3.8 GB).
In the agent harness the model runs 20-29 tok/s. A bf16 run of the same model measured 9.7 tok/s plain and 36.5 tok/s with DFlash 2 (4.43 accepted tokens per round, 137 target forwards, output identical). Agent-12 scores with DFlash 2 were 12/12 easy in 122 s and 7/8 hard in 401 s, about 3x faster than the Aug 18 run at the same accuracy. The Claude Code smoke test passed in 38 s. Gemma 4 31B Abliterated 4-bit ran at 26 tok/s and Gemma 4 E4B 4-bit failed the smoke test. On an M4 Pro 64 GB, Gemma 4 31B runs at ~13.5 tok/s and server-side prompt fixes cut hello latency from ~120 s to ~3-5 s.