llamaperf

Qwen3.8 125B (6B active) Flash-Next

on M5 Ultra 96GB · mlx-serve · 131,072 ctx

Tone: positive
Sep 23, 2026
Throughput
42.7 t/s gen · 828.2 t/s pp
Quant
4-bit-8bit mix (MLX)
KV cache
8-bit
System RAM
96 GB
VRAM reported
90 GB

Use cases

codingagenticlong-context

Summary

User reports Qwen3.8 Flash-Next on a base M5 Ultra 96GB (64-core GPU) at a median 828.2 t/s prefill and 42.7 t/s decode per stream, with a single-stream max of 3,328 t/s prefill and 152 t/s decode. Setup is a custom mlx-serve build with continuous batching at 4-way concurrency, a 4-bit/8-bit mixed MLX quant with MTP, 4 x 128k context (512k total) and 8-bit KV cache using 90GB of unified memory. Aggregate throughput was ~3,200 t/s prefill and ~170.8 t/s decode at 4-way concurrency. The run covered 112M tokens (109M prompt, 3M generated) with an 89% cache hit rate across ~1,700 sub-agent calls.