llamaperf

Qwen3.8 27B

on M3 Ultra 256GB · mlx-serve · 12,000 ctx

Oct 6, 2026
Throughput
84.1 t/s gen · 285-325 t/s pp
Quant
4-bit (MLX)
System RAM
256 GB

Summary

User benchmarks Qwen3.8-27B at 84.1 and 72.1 tok/s on production-shaped prompts and 87.0 tok/s at 12K context on a Mac Studio M3 Ultra 256GB. Setup is mlx-serve 26.9.6-dev with a 4-bit MTP pack and native MTP; prefill measured at roughly 285-325 tok/s, and concurrency gave no gain (1.01x at N=4). The same model class under Inco Splash reached 85.3-85.7 tok/s on short prompts and 65.7 tok/s at 32K, scaling 1.18x at N=2 and 1.27-1.29x at N=4 before going flat, with the engine reporting a maximum batch width of 4. The user notes the two engines' single-stream figures came from different prompts, so the 27B single-stream comparison is left undecided. Also measured: Qwen3.6-35B-A3B with mlx-dspark at accept length 4.5998 and 2.33x, versus MTPLX at accept 2.2294 and about 1.47x; Qwen3.8-Flash-Next with mlx-serve 26.9.6 and --mtp at 76.8 tok/s single stream (+32% over the same build without the flag) and 133.5 tok/s aggregate at N=8.