Qwen3.8 125B (6B active) Flash-Next
on AMD Radeon Pro W7800 48GB · llama.cpp · 131,072 ctx
Oct 5, 2026
Use cases
coding
Summary
User reports Qwen 3.8 Flash-Next at about 90 t/s on a Radeon Pro W7800 48GB with 64GB DDR5 system RAM.
Setup is llama.cpp with an IQ3_XXS GGUF quant, 128k context and Q8 KV cache, run with default settings.
The user is satisfied with the speed but the model fell into repetition loops after 3-4 minutes of reasoning on a 27KB JavaScript audit task, repeating the same steps dozens of times; they interrupted it twice and ask whether a reasoning budget setting is the cause.