llamaperf

Qwen3.8 125B (6B active) Flash-Next

on Unknown GPU · Strata · 524,288 ctx

Oct 5, 2026
Throughput
85.0 t/s gen
Quant
IQ2_XS (GGUF)
KV cache
q4_0

Use cases

long-context

Summary

User shares a tuning guide prompt for Strata, citing reference figures from a single RTX 4090 with 32 GB RAM running Qwen3.8-Flash-Next at IQ2_XS. Reference setup uses q4_0 KV cache at 262K native context, and yarn rope scaling with rope-scale 2, --max-context 524288 and a streamed KV cache (--kv-resident 32768) for 512K. A 477K-token prompt read in about 150 s and decoded at about 85 tok/s at that depth on that box. The guide is a template for other users to fill in their own hardware; the 4090 figures are starting points, not targets. 512K is marked experimental and requires a recall test at 400K+.