Qwen3 8B
on AMD RX 6700 XT · llama.cpp
Oct 5, 2026
Summary
User reports Qwen3-8B Q4_K_M at 851 t/s prompt and 61 t/s generation on an RX 6700 XT 12 GB.
Setup is llama.cpp with the AMD Flash Attention kernel, KV cache f16, measured with pp512 / tg128.
The project also lists Qwen3.6-35B-A3B Q4_K_S at 475 t/s prompt and 29 t/s generation with --n-cpu-moe 24, and gpt-oss-20B Q4_K_M at 1305 t/s prompt and 94 t/s generation with all experts in VRAM.