llamaperf

Qwen3.8 125B (6B active) Flash-Next

on NVIDIA RTX 5070 · Strata · 131,072 ctx

Tone: positive
Oct 5, 2026
Throughput
53.0 t/s gen · 1620.0 t/s pp
Quant
IQ3_S (GGUF)
System RAM
64 GB
VRAM reported
12 GB

Use cases

codingvisionagenticlong-context

Summary

User reports Qwen3.8-Flash-Next at 53 tokens/s generation and 1,620 tokens/s prompt processing on an RTX 5070 12GB with 64GB RAM. Setup is the Strata engine with IQ3_S quantization at 128K context; the model's experts are split across GPU, RAM and SSD. The page also lists Q2_0 at 94/2,650, IQ2_XS at 79/2,090, IQ3_XXS at 62/1,750 and Coder at 55/2,180 tokens/s on the same card, and RX 9070 XT 16GB figures of 60/1,160 (Q2_0), 52/1,110 (IQ2_XS) and 44/1,420 (Coder). Unsloth UD-Q4_K_XL is given as 7-8.5 tokens/s on a 64GB PC.