llamaperf
Sep 29, 2026
Throughput
33.5 t/s gen · 1210.3 t/s pp
Quant
Q4_K_M (GGUF)
System RAM
32 GB
VRAM reported
16 GB

Summary

User benchmarks Qwen 2.5-Coder 14B Instruct Q4_K_M at 33.51 t/s generation and 1210.29 t/s prompt processing on an AMD RX 9060 XT 16GB via Vulkan. Setup is llama.cpp build c4ae9a88f8 with -ngl 99 and -fa 1; the same model on ROCm reaches 30.79 t/s generation and 1211.32 t/s prompt processing. A Qwen 3.5 9B Q4_K_M run reaches 50.15 t/s generation and 1904.16 t/s prompt processing on Vulkan, and a 14B quant sweep gives about 32, 28 and 20 tok/s for Q4_K_M, Q5_K_M and Q8_0.