Qwen3.8 125B (6B active) Flash-Next
AMD RX 6900 XT 16GB · Strata · 64,000 ctx
- reported speed:
- 25.0 tokens/s generation
- quant:
- GSQ-RCO
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3.8-Flash-Next at 25 tok/s on an RX 6900 XT installed in a retired HP DL380p server with 172 GB DDR3 and 2x 10-core CPUs. Setup is Strata with llama.cpp and an HTTP router, GSQ-RCO quant at 64k context. The GPU is powered by an external desktop PSU; total draw is 250 W idle and about 450 W while inferring. User also reports Qwen3.8-27B with GSQ-RCO-IQ3_XXS and MTP at 192k context reaching 45 tok/s. The machine is a prototype; user plans a flexible riser and a proper GPU platform, and may add Nvidia P40s in the remaining slots someday.