Qwen3-Coder-Next
AMD Strix Halo 128GB · llama.cpp · 262,144 ctx
- reported speed:
- 36.8 tokens/s generation · 545.8 tokens/s prompt processing
- quant:
- UD-Q6_K_XL (GGUF)
- kv:
- f16
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
codingagenticlong-context
Post also mentions a 27B 8-bit XL model that was too slow to be workable, but the benchmarked/served model is Qwen3-Coder-Next UD-Q6_K_XL. Bench run with llama-benchy 0.4.1 API latency mode at -c 262144. tg32 peak 37.94 t/s.