Qwen3.8 125B (6B active) Flash-Next
2× NVIDIA RTX 3090 · Strata · 262,144 ctx
- reported speed:
- 43.9 tokens/s generation · 1817.0 tokens/s prompt processing
- quant:
- Q6_K_XL (GGUF)
- kv:
- int8
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3.8-Flash-Next at 43.9 tok/s generation and 1817 tok/s prompt processing on an RTX 3090 + RTX A4000 at 32k context. Setup is Strata with UD-Q6_K_XL GGUF and int8 KV cache, 262144 native context, speculative decoding with MTP, experts offloaded to CPU. At 200k context Strata Q6 does 48.8 tok/s versus llama.cpp's 10.6 tok/s, and TTFT drops from 595 s to 118 s. Q6 is 12-31% faster than Q8 and 19 GB smaller.