llamaperf
Oct 4, 2026
Throughput
154.5 t/s gen · 4183.0 t/s pp
Quant
mixed 4/8-bit (MLX)
KV cache
bf16
System RAM
256 GB

Summary

User benchmarks Qwen 3.8 Flash-Next (125B-A6B MoE) at 154.5 tok/s decode and 4,183 tok/s prefill on a single Mac Studio M5 Ultra 256 GB. Setup is mlx-serve 26.9.7-dev with the MLX-Serve mixed 4/8-bit pack (100 GiB), KV cache bf16, MTP speculative decoding, PLD and batched decode, at 262,144 context. Decode holds ~150 tok/s through 66k and is 114.4 tok/s at 256k; prefill stays flat at ~4.1-4.35k tok/s. A 4-stream burst gives 166.8 tok/s aggregate versus 119.1 alone. The user notes long-context recall was not verified because answers never cited the planted value.