Qwen3.8 125B (6B active) Flash-Next
on M5 Ultra 256GB · mlx-serve · 262,144 ctx
Oct 4, 2026
Summary
User benchmarks Qwen 3.8 Flash-Next (125B-A6B MoE) at 154.5 tok/s decode and 4,183 tok/s prefill on a single Mac Studio M5 Ultra 256 GB.
Setup is mlx-serve 26.9.7-dev with the MLX-Serve mixed 4/8-bit pack (100 GiB), KV cache bf16, MTP speculative decoding, PLD and batched decode, at 262,144 context.
Decode holds ~150 tok/s through 66k and is 114.4 tok/s at 256k; prefill stays flat at ~4.1-4.35k tok/s. A 4-stream burst gives 166.8 tok/s aggregate versus 119.1 alone. The user notes long-context recall was not verified because answers never cited the planted value.