Qwen3-Coder 30B (3B active)
AMD MI50 32GB · llama.cpp
- reported speed:
- 66.1 tokens/s generation
- quant:
- Q4_K_M (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User benchmarks Qwen3 Coder 30B A3B at 66.11 t/s generation on one AMD Radeon Instinct Mi50 32GB. Setup is llama.cpp build 128d522c (6686) with the ROCm backend, Q4_K_M quant (17.28 GiB), -ngl 99, 112 threads and --numa distribute on a dual Xeon 8480+ host. Other quants on the same card: Q5_0 65.15 t/s, Q6_K 62.49 t/s, Q8_0 64.20 t/s. BF16 (56.89 GiB) required two GPUs and ran 41.41 t/s. With 16K tokens of output the rate falls to about 20 t/s.