llamaperf

Qwen3.8 125B (6B active) Flash-Next

on RTX 2060 Laptop 6GB · Strata · 50,000 ctx

Tone: positive
Oct 3, 2026
Throughput
10.0 t/s gen
Quant
q2_0 (GGUF)
KV cache
Q4_0
System RAM
32 GB
VRAM reported
6 GB

Summary

User reports Qwen 3.8 Flash Next at 10 t/s decode and almost 100 t/s prefill on an RTX 2060 Laptop 6GB with 32 GB RAM. Setup is Strata with q2_0 quant and Q4_0 KV cache at 50k context, using draft-vocab en and 100 MiB VRAM reserve. User notes prefill is the bottleneck for large prompts, and that Qwen A35B A3B has around 5x faster prefill and allows 100K context without KV quantization.