llamaperf

Qwen3.8 125B (6B active) Flash-Next

on M5 128GB · oMLX

Tone: positive
Oct 6, 2026
Throughput
40.0 t/s gen
System RAM
128 GB

Use cases

codingagentic

Summary

User reports Qwen3.8-Flash-Next at 40 t/s on an M5 Mac with 128GB unified memory. Setup is oMLX runtime, model fits in 90-95GB of RAM. User notes MTPLX reaches 60 t/s with more tuning, and Qwen3.6 MoE on Splash runtime hits 120 t/s for most tasks except coding.