llamaperf

RWKV7

1 report

As of 8 Oct 2026, RWKV7 2.9B at 16-bit on the hardware it is most run on, with the median of plain runs (one device, one request, no speculative decoding, the whole model in its memory):

RWKV7 VRAM requirements by size and quant →

How does RWKV7 run on your hardware?

Pick your GPU or Mac, then read the speeds people reported for RWKV7 on it, or estimate memory fit and speed.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Run RWKV7 yourself? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance for RWKV7

Filter this model’s reports by setup →
Thin page (1 of 3 reports needed for indexing). Add yours.
reported speed:
48.7 tokens/s generation · 4315.7 tokens/s prompt processing
quant:
F16 (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports RWKV7 2.9B at 48.71 t/s generation and 4315.67 t/s prompt processing on an AMD Radeon Pro W7900. Setup is llama.cpp with F16 weights and ROCm backend, 99 GPU layers, 512-token prompt and 128-token generation. Q8_0 on ROCm reaches 58.59 t/s generation and 4033.24 t/s prompt; Vulkan backend gives 39.49 t/s (F16) and 45.21 t/s (Q8_0) generation.

Oct 8, 2026

Get a weekly email of new RWKV7 reports on any GPU.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
RWKV7 2.9B
AMD Radeon Pro W7900
F16
llama.cpp
Not reported48.7 tokens/s