Qwen3.8 125B (6B active) Flash-Next
on M5 Ultra 256GB · oMLX · 1,000,000 ctx
Oct 3, 2026
Use cases
codinglong-context
Summary
User reports Qwen3.8-Flash-Next 8-bit at 29 t/s decode on a 1M-token prompt on an M5 Ultra 256GB Mac Studio.
Setup is oMLX 0.7.0 with oQ8e weights and YaRN x4 position scaling; the 1M-token first read took 5.3 minutes and the next turn 6.7 s, finding all three hidden codes.
At 250k context the model prefills at 4,235 t/s and decodes at 63 t/s; the user notes YaRN stretches a 262k-trained model and that a needle test does not prove reasoning at 1M.