llamaperf

Qwen3.8 125B (6B active) Flash-Next

on AMD Radeon Pro W7800 48GB · llama.cpp · 131,072 ctx

Tone: mixed
Oct 5, 2026
Throughput
90.0 t/s gen
Quant
IQ3_XXS (GGUF)
KV cache
Q8
System RAM
64 GB
VRAM reported
48 GB

Use cases

coding

Summary

User reports Qwen 3.8 Flash-Next at about 90 t/s on a Radeon Pro W7800 48GB with 64GB DDR5 system RAM. Setup is llama.cpp with an IQ3_XXS GGUF quant, 128k context and Q8 KV cache, run with default settings. The user is satisfied with the speed but the model fell into repetition loops after 3-4 minutes of reasoning on a 27KB JavaScript audit task, repeating the same steps dozens of times; they interrupted it twice and ask whether a reasoning budget setting is the cause.