Qwen3.8 125B (6B active) Flash-Next
AMD MI50 16GB · llama.cpp · 131,072 ctx
- reported speed:
- 16.9 tokens/s generation
- quant:
- IQ4 (GGUF)
- kv:
- Q8
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3.8 Flash-Next at 16.9 t/s on a single MI50 16GB with the gfx906-16gb-expert-pool llama.cpp fork. Setup uses the moe-expert-pool branch with IQ4 GGUF weights, 128K context and Q8 KV cache; the expert cache profile of 66 admits 144 pools at roughly 69% hit rate and peaks at 15.24 GiB VRAM. A stock CPU MoE path without the expert cache reached 11.76 t/s, and up to 19.8 t/s was observed once warm.