Qwen3.8 125B (6B active) Flash-Next
on M4 Max 64GB · mlx-serve · 65,536 ctx
Sep 27, 2026
Use cases
codingagenticvisionlong-context
Summary
User reports Qwen3.8-Flash-Next at 52.6 tok/s decode and 762 tok/s prefill on an M4 Max 64GB.
Setup is mlx-serve with an iQ-MLX 3.3bpw pack, 64k context, 52 GB resident; the 32 GB n-gram table is mmapped and not resident.
The same run measured the mixed-4-8bit pack at 55.5 tok/s decode, 754 tok/s prefill and 278 ms TTFT; MTP is in the pack but default-off and not re-measured.