llamaperf

Qwen3 8B

on AMD RX 6700 XT · llama.cpp

Oct 5, 2026
Throughput
61.0 t/s gen · 851.0 t/s pp
Quant
Q4_K_M (GGUF)
KV cache
f16
VRAM reported
12 GB

Summary

User reports Qwen3-8B Q4_K_M at 851 t/s prompt and 61 t/s generation on an RX 6700 XT 12 GB. Setup is llama.cpp with the AMD Flash Attention kernel, KV cache f16, measured with pp512 / tg128. The project also lists Qwen3.6-35B-A3B Q4_K_S at 475 t/s prompt and 29 t/s generation with --n-cpu-moe 24, and gpt-oss-20B Q4_K_M at 1305 t/s prompt and 94 t/s generation with all experts in VRAM.