llamaperf

Qwen3.8 125B (6B active) Flash-Next

on M5 Ultra 256GB · oMLX · 1,000,000 ctx

Tone: positive
Oct 3, 2026
Throughput
29.0 t/s gen
Quant
oQ8e (MLX)
System RAM
256 GB

Use cases

codinglong-context

Summary

User reports Qwen3.8-Flash-Next 8-bit at 29 t/s decode on a 1M-token prompt on an M5 Ultra 256GB Mac Studio. Setup is oMLX 0.7.0 with oQ8e weights and YaRN x4 position scaling; the 1M-token first read took 5.3 minutes and the next turn 6.7 s, finding all three hidden codes. At 250k context the model prefills at 4,235 t/s and decodes at 63 t/s; the user notes YaRN stretches a 262k-trained model and that a needle test does not prove reasoning at 1M.