Qwen3.8 125B (6B active) Flash-Next
on M3 Ultra 96GB · ds4 · 262,144 ctx
Oct 3, 2026
Use cases
coding
Summary
User reports Qwen3.8 Flash Next Q4 running through ds4 on an M3 Ultra 96GB with the full 262K context configured, delivering 55–60 tok/s decode and 667 tok/s prefill.
About 80GB is used for model, KV cache and buffers, while the model's 95GB n-gram table streams from SSD.
The user runs it daily through a codex harness as a headless box.