RWKV7 2.9B
AMD Radeon Pro W7900 · llama.cpp
- reported speed:
- 48.7 tokens/s generation · 4315.7 tokens/s prompt processing
- quant:
- F16 (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports RWKV7 2.9B at 48.71 t/s generation and 4315.67 t/s prompt processing on an AMD Radeon Pro W7900. Setup is llama.cpp with F16 weights and ROCm backend, 99 GPU layers, 512-token prompt and 128-token generation. Q8_0 on ROCm reaches 58.59 t/s generation and 4033.24 t/s prompt; Vulkan backend gives 39.49 t/s (F16) and 45.21 t/s (Q8_0) generation.