Qwen3.8 125B (6B active) Flash-Next
on M4 Max 128GB · mlx-serve · 262,144 ctx
Oct 7, 2026
Use cases
codinglong-contextvision
Summary
User reports Qwen3.8 Flash-Next at 73.8 tok/s on an M4 Max with 128 GB unified memory.
Setup is mlx-serve with Q4 affine routed experts, BF16 core and verification head, native MTP at depth 4, and the full 262,144-token context; the 51.2 GB n-gram lookup table streams from SSD.
Matched warm workloads measured 71-74 tok/s for code and technical prose, with serial decode at 33.2 tok/s and a 2.22x MTP speedup. Long-context prefill reached 280.7 prompt tok/s at 248,445 tokens.