Qwen3.8 125B (6B active) Flash-Next
on M5 Ultra 256GB · oMLX · 262,144 ctx
Sep 30, 2026
Summary
User compares Qwen3.8-Flash-Next against Laguna-S-2.1 on a Mac Studio M5 Ultra 256GB, both capped at 262K context with thinking on and unique content per run.
Qwen runs under oMLX with an oQ8e quant and MTP speculation, holding roughly 4,200 tok/s prefill and 59-74 tok/s decode across sizes. Laguna runs under LM Studio with an 8-bit quant, prefill degrading superlinearly from 10.3s at 8K to 455.4s at 200K, and decode falling from 68 to 34 tok/s with no speculation.
Quality was a draw at 4/4 each on four script-verified problems. At 200K, prefill is 94% of total time on both. The user notes an earlier run was invalidated by shared prefixes letting the KV cache carry over.