Qwen3.8 125B (6B active) Flash-Next
on M5 Ultra 256GB · oMLX · 16,384 ctx
Sep 29, 2026
Use cases
agentic
Summary
User reports Qwen3.8-Flash-Next at 108 t/s generation and 2,887 t/s prompt processing on an M5 Ultra 256GB Mac Studio.
Setup is oMLX 0.7.0.dev2 with MLX backend, oQ4e 4-bit quant, multi-token prediction depth 3, 16K prompt.
M3 Ultra 512GB comparison: 70 t/s generation, 1,143 t/s prompt. At 256K context, M5 prompt processing was 2,544 t/s. Qwen3.8-27B on M5 Ultra: 48 t/s generation, 1,701 t/s prompt; RTX 5090 PC in LM Studio: 59 t/s generation, 3,031 t/s prompt.