llamaperf

AMD RX 6700 XT

AMD · 12GB · 3 reports

As of 7 Oct 2026, the models most run on the AMD RX 6700 XT, with the median of plain runs (one device, one request, no speculative decoding):

Engines people use on it: llama.cpp 3

Run models on your AMD RX 6700 XT? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the AMD RX 6700 XT

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 12 GB of VRAM.

Gemma 4 12B

AMD RX 6700 XT · llama.cpp · 8,192 ctx

Tone: mixed
reported speed:
34.6 tokens/s generation · 653.9 tokens/s prompt processing
quant:
IQ4_NL (GGUF)
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

long-context

User benchmarks Gemma 4 12B (IQ4_NL, 6.24 GiB) on an AMD RX 6700 XT under llama.cpp, comparing the ROCm and Vulkan backends at 8192-token prefill and 512-token generation with q8_0 KV cache and flash-attention on. ROCm averages 653.9 t/s prefill and 34.60 t/s decode over 3 runs; Vulkan averages 354.4 t/s prefill and 40.92 t/s decode. ROCm is 84.5% faster on prefill but 15.4% slower on decode, giving a net wall-clock win of about 23% for a full 8192-prefill plus 512-generate cycle, with a crossover near 1760 prompt tokens. ROCm required two workarounds on gfx1031: building for gfx1030 with HSA_OVERRIDE_GFX_VERSION=10.3.0, and patching a flash-attention assert in fattn-common.cuh. A separate TOP_K sampler gap on ROCm is noted as under investigation.

Oct 6, 2026
reported speed:
61.0 tokens/s generation · 851.0 tokens/s prompt processing
quant:
Q4_K_M (GGUF)
kv:
f16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3-8B Q4_K_M at 851 t/s prompt and 61 t/s generation on an RX 6700 XT 12 GB. Setup is llama.cpp with the AMD Flash Attention kernel, KV cache f16, measured with pp512 / tg128. The project also lists Qwen3.6-35B-A3B Q4_K_S at 475 t/s prompt and 29 t/s generation with --n-cpu-moe 24, and gpt-oss-20B Q4_K_M at 1305 t/s prompt and 94 t/s generation with all experts in VRAM.

Oct 5, 2026

Qwen3.8 27B

2× AMD RX 6700 XT · llama.cpp · 144,000 ctx

Tone: positive
reported speed:
18.0 tokens/s generation · 150.0 tokens/s prompt processing
quant:
UD-IQ3_XXS (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticlong-context

User reports Qwen3.8 27B at 18 t/s generation and 150 t/s prompt processing on a 2007 Dell Precision T5400 with dual RX 6700 XT / RX 6700 (22GB total VRAM). Setup is llama.cpp with UD-IQ3_XXS quant at 144k context, using MTP q4_0 speculative decoding, on 24GB DDR2 and dual Xeon X5460. User compares five systems and argues older dual-Xeon dual-GPU setups beat a 2025 HP Omen with RTX 5070 in context length and speed, concluding DDR5 is not worth the cost for agentic tasks.

Sep 28, 2026

Get a weekly email of new AMD RX 6700 XT reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Gemma 4 12B
AMD RX 6700 XT
IQ4_NL
llama.cpp
8,19234.6 tokens/s
Qwen3 8B
AMD RX 6700 XT
Q4_K_M
llama.cpp
Not reported61.0 tokens/s
Qwen3.8 27B
2× AMD RX 6700 XT
UD-IQ3_XXS
llama.cpp
144,00018.0 tokens/s