llamaperf

Qwen3.8 125B (6B active) Flash-Next

on 2× NVIDIA RTX 5080 · Strata · 65,536 ctx

Tone: positive
Oct 3, 2026
Throughput
110.6 t/s gen · 2773.0 t/s pp
Quant
UD-Q4_K_XL (GGUF)
KV cache
INT8
VRAM reported
16 GB

Use cases

coding

Summary

User reports Qwen3.8-Flash-Next at 110.6 t/s generation and 2773 t/s prefill on two RTX 5080s with a Ryzen 9 9900X, using a custom Strata fork. Setup is Strata with UD-Q4_K_XL, 64K context, INT8 K/V, 8K prompt chunk, MTP on and expert adaptation on. Upstream v0.1.38 reached 78.1 t/s generation and 2016 t/s prefill with layer split, and 70.3 t/s generation and 1522 t/s prefill with the host peer tier. With static caches custom was about 26% ahead of upstream layer split (55 vs 44 t/s). Letting the second GPU compute experts during prompt processing doubled custom's 32K prefill (1366 to 2731 t/s), and a 16K chunk brought it to 3724 t/s. Putting the expert tier on the x8 PCIe card improved 4K prefill by 58% and 32K by 34%. The saved answer came from the custom fork, so it may favor its drafts.