llamaperf
Sep 30, 2026
Throughput
66.1 t/s gen
Quant
Q4_K_M (GGUF)
VRAM reported
32 GB

Use cases

coding

Summary

User benchmarks Qwen3 Coder 30B A3B at 66.11 t/s generation on one AMD Radeon Instinct Mi50 32GB. Setup is llama.cpp build 128d522c (6686) with the ROCm backend, Q4_K_M quant (17.28 GiB), -ngl 99, 112 threads and --numa distribute on a dual Xeon 8480+ host. Other quants on the same card: Q5_0 65.15 t/s, Q6_K 62.49 t/s, Q8_0 64.20 t/s. BF16 (56.89 GiB) required two GPUs and ran 41.41 t/s. With 16K tokens of output the rate falls to about 20 t/s.