llamaperf

Qwen3.8 125B (6B active) Flash-Next

on M3 Ultra · oMLX · 262,144 ctx

Oct 6, 2026
Throughput
29.1 t/s gen
Quant
oQ3 (MLX)

Summary

User reports Qwen3.8-Flash-Next at 29.0822 tokens/s median generation on an Apple M3 Studio. Setup is oMLX 0.6.3rc3 with MLX-VLM 0.6.3 and MLX 0.32.0, oQ3 mixed-precision MLX safetensors (3-bit affine base, 92.505 GB), native MTP enabled, 262,144-token configured context, greedy decoding over three 512-token runs. With native MTP disabled the same setup reached 26.5352 tokens/s, a 9.60% improvement; MTP drafted 847 tokens and accepted 583 (68.83% acceptance), and both modes produced matching output hashes.