llamaperf
Oct 8, 2026
Throughput
48.7 t/s gen · 4315.7 t/s pp
Quant
F16 (GGUF)
System RAM
64 GB

Summary

User reports RWKV7 2.9B at 48.71 t/s generation and 4315.67 t/s prompt processing on an AMD Radeon Pro W7900. Setup is llama.cpp with F16 weights and ROCm backend, 99 GPU layers, 512-token prompt and 128-token generation. Q8_0 on ROCm reaches 58.59 t/s generation and 4033.24 t/s prompt; Vulkan backend gives 39.49 t/s (F16) and 45.21 t/s (Q8_0) generation.