llamaperf

Qwen3.8 27B

on M2 Max 96GB · mlx-serve

Tone: mixed
Sep 23, 2026
Throughput
20.1 t/s gen
Quant
4bit (MLX)
System RAM
96 GB

Summary

User benchmarks Qwen3.8 27B 4-bit on an M2 Max 96GB, comparing mlx-serve with MTP speculative decoding against LM Studio. With the MTP head correctly loaded, mlx-serve reaches 20.1 t/s decode on long prompts and 13.0 t/s on short, versus 11.2 and 8.0 t/s without speculation. Setup is mlx-serve 26.9.5 with MLX 0.32.2 on macOS 26.6.2, using the mlx-community Qwen3.8-27B-MTP-4bit sidecar. Prefill measured 49/145 tok/s with MTP loaded and 47/112 tok/s without. LM Studio on the same weights reached 11.8 t/s decode short and 11.6 t/s long, with 59/101 tok/s prefill. The user reports that mlx-serve silently falls back to Prompt Lookup Decoding when the MTP sidecar tensors lack the required mtp. prefix, and that MoE MTP heads fail to load entirely, so Qwen3.6-35B-A3B got no speedup. Numbers are client-side medians of 5 runs with unique prefixes to avoid prefix caching.