llamaperf
Sep 27, 2026
Throughput
706.0 t/s pp
Quant
Q4_K_M (GGUF)
System RAM
64 GB
VRAM reported
16 GB

Summary

User reports Qwen3.6-35B-A3B at 706.0 t/s prompt processing on an AMD Radeon RX 9070 XT 16GB under ROCm. Setup is llama.cpp with Q4_K_M GGUF, flash attention enabled, -ncmoe 40 and -ub 512, on a Ryzen 9 5950X with 64 GB DDR4-3200. The same configuration under Vulkan reached 333.4 t/s prompt and 29.3 t/s generation; the user retired Vulkan in favour of ROCm.