Qwen3.8 125B (6B active) Flash-Next
AMD Radeon Pro W7800 48GB · llama.cpp · 131,072 ctx
- reported speed:
- 90.0 tokens/s generation
- quant:
- IQ3_XXS (GGUF)
- kv:
- Q8
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen 3.8 Flash-Next at about 90 t/s on a Radeon Pro W7800 48GB with 64GB DDR5 system RAM. Setup is llama.cpp with an IQ3_XXS GGUF quant, 128k context and Q8 KV cache, run with default settings. The user is satisfied with the speed but the model fell into repetition loops after 3-4 minutes of reasoning on a 27KB JavaScript audit task, repeating the same steps dozens of times; they interrupted it twice and ask whether a reasoning budget setting is the cause.