llamaperf

Qwen3-Coder-Next

1 report

Thin page (1 of 3 reports needed for indexing). Add yours.

Qwen3-Coder-Next

AMD Strix Halo 128GB · llama.cpp · 262,144 ctx

Tone: mixed
reported speed:
36.8 tokens/s generation · 545.8 tokens/s prompt processing
quant:
UD-Q6_K_XL (GGUF)
kv:
f16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagenticlong-context

Post also mentions a 27B 8-bit XL model that was too slow to be workable, but the benchmarked/served model is Qwen3-Coder-Next UD-Q6_K_XL. Bench run with llama-benchy 0.4.1 API latency mode at -c 262144. tg32 peak 37.94 t/s.