llamaperf

Qwen3.8 125B (6B active) Flash-Next

on M4 Max 128GB · mlx-serve · 262,144 ctx

Tone: positive
Oct 7, 2026
Throughput
73.8 t/s gen
Quant
Q4 (MLX)
System RAM
128 GB

Use cases

codinglong-contextvision

Summary

User reports Qwen3.8 Flash-Next at 73.8 tok/s on an M4 Max with 128 GB unified memory. Setup is mlx-serve with Q4 affine routed experts, BF16 core and verification head, native MTP at depth 4, and the full 262,144-token context; the 51.2 GB n-gram lookup table streams from SSD. Matched warm workloads measured 71-74 tok/s for code and technical prose, with serial decode at 33.2 tok/s and a 2.22x MTP speedup. Long-context prefill reached 280.7 prompt tok/s at 248,445 tokens.