Qwen3.8 125B (6B active) Flash-Next
on Unknown GPU · Strata · 524,288 ctx
Oct 5, 2026
Use cases
long-context
Summary
User shares a tuning guide prompt for Strata, citing reference figures from a single RTX 4090 with 32 GB RAM running Qwen3.8-Flash-Next at IQ2_XS.
Reference setup uses q4_0 KV cache at 262K native context, and yarn rope scaling with rope-scale 2, --max-context 524288 and a streamed KV cache (--kv-resident 32768) for 512K. A 477K-token prompt read in about 150 s and decoded at about 85 tok/s at that depth on that box.
The guide is a template for other users to fill in their own hardware; the 4090 figures are starting points, not targets. 512K is marked experimental and requires a recall test at 400K+.