llamaperf

Qwen3.8 27B

on M5 Max 128GB · mlx-vlm

Tone: positive
Oct 3, 2026
Throughput
47.0 t/s gen · 854.0 t/s pp
Quant
4bit (MLX)
System RAM
128 GB

Use cases

codingagentic

Summary

User reports Qwen3.8-27B at 47.0 t/s generation and 854 t/s prompt processing on an Apple M5 Max with 128 GB unified memory. Setup is mlx-vlm with the 4-bit MLX quantisation and MTP speculative decoding enabled, 15.5 GB resident. The 8-bit quant reaches 28.7 t/s generation and 883 t/s prompt processing at 28.1 GB resident. MTP speculative decoding is reported as worth 2.2x (15.5 t/s to 34.1 t/s at 8-bit). Decode falls to 36.1 t/s at 8K context, and time to first token climbs to about 14 s at that depth.