llamaperf

Qwen3.8 125B (6B active) Flash-Next

on NVIDIA RTX 5070 · Strata · 131,072 ctx

Tone: positive
Sep 24, 2026
Throughput
44.8 t/s gen · 414.0 t/s pp
Quant
IQ3_XXS (GGUF)
System RAM
64 GB
VRAM reported
12 GB

Summary

User reports Qwen3.8-Flash-Next at 44.8 t/s generation and 414 t/s prompt processing on a 12GB RTX 5070 at 128K context. Setup is a custom Strata inference engine with the IQ3_XXS GGUF quant, 64GB DDR5-5600 and a Ryzen 5 7600 on Windows. The engine is CUDA-only and requires 47GB minimum in RAM+VRAM for this quant. A Q2_0 quant reached 65.1 t/s generation and 543 t/s prompt processing, and IQ2_XS reached 52.0 t/s generation and 472 t/s prompt processing. The user previously measured 15 t/s output and 100-120 t/s prompt processing with llama.cpp on the same hardware.