RWKV7 2.9B
on AMD Radeon Pro W7900 · llama.cpp
Oct 8, 2026
Summary
User reports RWKV7 2.9B at 48.71 t/s generation and 4315.67 t/s prompt processing on an AMD Radeon Pro W7900.
Setup is llama.cpp with F16 weights and ROCm backend, 99 GPU layers, 512-token prompt and 128-token generation.
Q8_0 on ROCm reaches 58.59 t/s generation and 4033.24 t/s prompt; Vulkan backend gives 39.49 t/s (F16) and 45.21 t/s (Q8_0) generation.