Qwen3.8 125B (6B active) Flash-Next
on M5 Ultra 96GB · mlx-serve · 131,072 ctx
Sep 23, 2026
Use cases
codingagenticlong-context
Summary
User reports Qwen3.8 Flash-Next on a base M5 Ultra 96GB (64-core GPU) at a median 828.2 t/s prefill and 42.7 t/s decode per stream, with a single-stream max of 3,328 t/s prefill and 152 t/s decode.
Setup is a custom mlx-serve build with continuous batching at 4-way concurrency, a 4-bit/8-bit mixed MLX quant with MTP, 4 x 128k context (512k total) and 8-bit KV cache using 90GB of unified memory.
Aggregate throughput was ~3,200 t/s prefill and ~170.8 t/s decode at 4-way concurrency. The run covered 112M tokens (109M prompt, 3M generated) with an 89% cache hit rate across ~1,700 sub-agent calls.