llamaperf

Qwen3.8 125B (6B active) Flash-Next

on M5 Ultra 256GB · oMLX · 262,144 ctx

Tone: positive
Sep 30, 2026
Throughput
59-74 t/s gen
Quant
oQ8e (MLX)
System RAM
256 GB

Summary

User compares Qwen3.8-Flash-Next against Laguna-S-2.1 on a Mac Studio M5 Ultra 256GB, both capped at 262K context with thinking on and unique content per run. Qwen runs under oMLX with an oQ8e quant and MTP speculation, holding roughly 4,200 tok/s prefill and 59-74 tok/s decode across sizes. Laguna runs under LM Studio with an 8-bit quant, prefill degrading superlinearly from 10.3s at 8K to 455.4s at 200K, and decode falling from 68 to 34 tok/s with no speculation. Quality was a draw at 4/4 each on four script-verified problems. At 200K, prefill is 94% of total time on both. The user notes an earlier run was invalidated by shared prefixes letting the KV cache carry over.