llamaperf

Qwen3.8 125B (6B active) Flash-Next

on M3 Ultra 96GB · ds4 · 262,144 ctx

Tone: positive
Oct 3, 2026
Throughput
55-60 t/s gen · 667.0 t/s pp
Quant
Q4 (GGUF)
System RAM
96 GB

Use cases

coding

Summary

User reports Qwen3.8 Flash Next Q4 running through ds4 on an M3 Ultra 96GB with the full 262K context configured, delivering 55–60 tok/s decode and 667 tok/s prefill. About 80GB is used for model, KV cache and buffers, while the model's 95GB n-gram table streams from SSD. The user runs it daily through a codex harness as a headless box.