llamaperf

Qwen3.8 125B (6B active) Flash-Next

on NVIDIA RTX 3090 · Strata · 256,000 ctx

Tone: positive
Oct 3, 2026
Throughput
38-61 t/s gen · 1650.0 t/s pp
Quant
UD-Q3_K_XL (GGUF)
KV cache
fp16
System RAM
128 GB
VRAM reported
24 GB

Use cases

codingtool-uselong-context

Summary

User reports Qwen3.8 Flash-Next running on an RTX 3090 24GB with 128GB system RAM, reaching about 1650 t/s prompt processing and 38 to 61 t/s generation depending on context under Strata. Setup is Strata with Unsloth UD-Q3_K_XL quant, fp16 KV cache, 256k context, speculative decoding with 4 draft tokens, and an expert cache of 6517 slots (~14GB). The same model on llama.cpp master gave up to 700 t/s PP and 23 t/s TG. Generation is a range: 38 t/s at 182k context and about 61 t/s at short context. The user notes the quant Strata recommends by default was faster but produced minor errors and lower quality, so Unsloth's was chosen for accuracy.