llamaperf

Qwen3.8 125B (6B active) Flash-Next

on AMD Radeon Pro W7900 · llama-halo-hybrid · 128,000 ctx

Tone: mixed
Oct 6, 2026
Throughput
63.5 t/s gen · 1119.0 t/s pp
Quant
UD-Q4_K_XL (GGUF)
System RAM
128 GB
VRAM reported
48 GB

Use cases

codingcreative-writing

Summary

User reports Qwen3.8-Flash-Next at 63.5 t/s decode (code) and 1119 t/s prefill at 2.6K context on a Radeon Pro W7900 eGPU plus Strix Halo iGPU hybrid setup. Setup is llama-halo-hybrid fork with UD-Q4_K_XL quant at 128K context, MTP head enabled, dense layers on the W7900 and routed experts on the iGPU. Hybrid beats gufo iGPU-only (44.7 t/s decode code, 903 t/s prefill at 2.6K) and leaves 43-52 GiB RAM free versus 14-30. GLM-5.3-Flash UD-IQ3_XXS does 25 t/s (31.6 with MTP) and DS-V4-Flash UD-IQ3_XXS does 22.5 t/s. W7900 runs 100-104 C junction on long prefills with a locked 241 W power cap. One run each, no quality testing.