Qwen3.8 125B (6B active) Flash-Next
on AMD Radeon Pro W7900 · llama-halo-hybrid · 128,000 ctx
Oct 6, 2026
Use cases
codingcreative-writing
Summary
User reports Qwen3.8-Flash-Next at 63.5 t/s decode (code) and 1119 t/s prefill at 2.6K context on a Radeon Pro W7900 eGPU plus Strix Halo iGPU hybrid setup.
Setup is llama-halo-hybrid fork with UD-Q4_K_XL quant at 128K context, MTP head enabled, dense layers on the W7900 and routed experts on the iGPU.
Hybrid beats gufo iGPU-only (44.7 t/s decode code, 903 t/s prefill at 2.6K) and leaves 43-52 GiB RAM free versus 14-30. GLM-5.3-Flash UD-IQ3_XXS does 25 t/s (31.6 with MTP) and DS-V4-Flash UD-IQ3_XXS does 22.5 t/s. W7900 runs 100-104 C junction on long prefills with a locked 241 W power cap. One run each, no quality testing.