llamaperf

Qwen3.8 125B (6B active) Flash-Next

on 2× NVIDIA RTX 3090 · Strata · 262,144 ctx

Tone: positive
Oct 7, 2026
Throughput
43.9 t/s gen · 1817.0 t/s pp
Quant
Q6_K_XL (GGUF)
KV cache
int8
System RAM
278 GB

Summary

User reports Qwen3.8-Flash-Next at 43.9 tok/s generation and 1817 tok/s prompt processing on an RTX 3090 + RTX A4000 at 32k context. Setup is Strata with UD-Q6_K_XL GGUF and int8 KV cache, 262144 native context, speculative decoding with MTP, experts offloaded to CPU. At 200k context Strata Q6 does 48.8 tok/s versus llama.cpp's 10.6 tok/s, and TTFT drops from 595 s to 118 s. Q6 is 12-31% faster than Q8 and 19 GB smaller.