llamaperf

AMD RX 6600 XT

AMD · 8GB · 2 reports

Engines people use on it: llama.cpp 2

Run models on your AMD RX 6600 XT? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the AMD RX 6600 XT

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 8 GB of VRAM.

This page is thin (2 of 3 reports needed for indexing). Help fill it in.

Unknown family

AMD RX 6600 XT · llama.cpp · 130,416 ctx

reported speed:
89.7 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports prefill speed on an RX 6600 XT 8GB that does not fall steadily with context, using llama.cpp Vulkan with every layer offloaded: 935 t/s at about 1K tokens down to 318 t/s at 8K, then 472 t/s at 16K, and 89.7 t/s at 130K against 76.7 t/s at 65K. The model is not named. User says the spread across runs was under 2%.

Sep 26, 2026

Qwen3.6 35B (3B active)

AMD RX 6600 XT · llama.cpp · 65,536 ctx

reported speed:
30.0 tokens/s generation
quant:
UD_Q4_K_XL (GGUF)
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agentic

User asks what performance P100 owners get, having ordered one for $80, and reports their current baseline on an RX 6600 XT with 32GB DDR4 3600. Current setup runs unsloth Qwen3.6 35B-A3B UD_Q4_K_XL in llama.cpp with MTP and --cpu-moe at 64k full-precision context, giving 30 t/s generation and 48-50 t/s decode, with prefill up to 800 t/s at 0 context and 700 t/s at 10k. User hopes the P100 can match the decode numbers and plans to tune -b and -ub for its higher core count; the card will go into a dedicated inference machine.

Sep 23, 2026

Get a weekly email of new AMD RX 6600 XT reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Qwen3.6 35B (3B active)
AMD RX 6600 XT
UD_Q4_K_XL
llama.cpp
65,53630.0 tokens/s