Qwen3.8 125B (6B active) Flash-Next
on AMD RX 6900 XT 16GB · Strata · 64,000 ctx
Oct 6, 2026
Summary
User reports Qwen3.8-Flash-Next at 25 tok/s on an RX 6900 XT installed in a retired HP DL380p server with 172 GB DDR3 and 2x 10-core CPUs.
Setup is Strata with llama.cpp and an HTTP router, GSQ-RCO quant at 64k context. The GPU is powered by an external desktop PSU; total draw is 250 W idle and about 450 W while inferring.
User also reports Qwen3.8-27B with GSQ-RCO-IQ3_XXS and MTP at 192k context reaching 45 tok/s. The machine is a prototype; user plans a flexible riser and a proper GPU platform, and may add Nvidia P40s in the remaining slots someday.