Qwen3.8 27B
on M5 Max 128GB · mlx-vlm
Oct 3, 2026
Use cases
codingagentic
Summary
User reports Qwen3.8-27B at 47.0 t/s generation and 854 t/s prompt processing on an Apple M5 Max with 128 GB unified memory.
Setup is mlx-vlm with the 4-bit MLX quantisation and MTP speculative decoding enabled, 15.5 GB resident. The 8-bit quant reaches 28.7 t/s generation and 883 t/s prompt processing at 28.1 GB resident.
MTP speculative decoding is reported as worth 2.2x (15.5 t/s to 34.1 t/s at 8-bit). Decode falls to 36.1 t/s at 8K context, and time to first token climbs to about 14 s at that depth.