llamaperf

Qwen3.8 125B (6B active) Flash-Next

on M5 Ultra 256GB · oMLX · 16,384 ctx

Tone: positive
Sep 29, 2026
Throughput
108.0 t/s gen · 2887.0 t/s pp
Quant
oQ4e (MLX)
System RAM
256 GB

Use cases

agentic

Summary

User reports Qwen3.8-Flash-Next at 108 t/s generation and 2,887 t/s prompt processing on an M5 Ultra 256GB Mac Studio. Setup is oMLX 0.7.0.dev2 with MLX backend, oQ4e 4-bit quant, multi-token prediction depth 3, 16K prompt. M3 Ultra 512GB comparison: 70 t/s generation, 1,143 t/s prompt. At 256K context, M5 prompt processing was 2,544 t/s. Qwen3.8-27B on M5 Ultra: 48 t/s generation, 1,701 t/s prompt; RTX 5090 PC in LM Studio: 59 t/s generation, 3,031 t/s prompt.