llamaperf

AMD RX 9060 XT 16GB

AMD · 16GB · 5 reports

As of 7 Oct 2026, the models most run on the AMD RX 9060 XT 16GB, with the median of plain runs (one device, one request, no speculative decoding):

Engines people use on it: llama.cpp 3

Run models on your AMD RX 9060 XT 16GB? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the AMD RX 9060 XT 16GB

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 16 GB of VRAM.

reported speed:
66.0 tokens/s generation
quant:
UD-Q4_K_XL (GGUF)
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.6-35B-A3B at 66.04 tokens/s on an AMD Radeon RX 9060 XT 16 GB with 32 GB of system RAM. Setup is llama.cpp (Vulkan backend) with UD-Q4_K_XL GGUF weights, q8_0 KV cache, 40k context, and MTP speculative decoding drafting up to 3 tokens per step. The 35B MoE model does not fit in 16 GB of VRAM, so the expert weights of the first 20 layers run on the CPU. Run on a Ryzen 7 5800X3D desktop under SteamOS, single slot, with flash attention and prefix cache reuse enabled.

Oct 6, 2026
reported speed:
30.0 tokens/s generation
quant:
IQ3_S (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User asks whether Qwen3.8 Flash-Next can run on an RX 9060 XT 16GB with 32GB DDR5-6000. They currently run Swift-1.5-Qwen3.8 at IQ3_S at 30 t/s on the same machine.

Oct 3, 2026
reported speed:
50.0 tokens/s generation · 850.0 tokens/s prompt processing
quant:
IQ3_S (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports 850 t/s prefill and 50 t/s generation with MTP at temperature 0 on an AMD RX 9060 XT 16GB, running Qwen3.8-27B-GSQ-RCO IQ3_S in a custom inference engine. MTP is used for the generation figure; the user notes MTP with temperature above 0 is not developed yet. User asks for ideas to systematically test the engine, having already run KL divergence against BF12 on CPU, bit correctness checks, needle-in-a-haystack at 25/50/75/90% key positions with distractors (also against llama.cpp), and HumanEval (92/93 passed so far).

Oct 3, 2026
reported speed:
33.5 tokens/s generation · 1210.3 tokens/s prompt processing
quant:
Q4_K_M (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks Qwen 2.5-Coder 14B Instruct Q4_K_M at 33.51 t/s generation and 1210.29 t/s prompt processing on an AMD RX 9060 XT 16GB via Vulkan. Setup is llama.cpp build c4ae9a88f8 with -ngl 99 and -fa 1; the same model on ROCm reaches 30.79 t/s generation and 1211.32 t/s prompt processing. A Qwen 3.5 9B Q4_K_M run reaches 50.15 t/s generation and 1904.16 t/s prompt processing on Vulkan, and a 14B quant sweep gives about 32, 28 and 20 tok/s for Q4_K_M, Q5_K_M and Q8_0.

Sep 29, 2026

Ornith1.5 9B

AMD RX 9060 XT 16GB · llama.cpp · 262,144 ctx

Tone: positive
reported speed:
36.0 tokens/s generation · 950.0 tokens/s prompt processing
quant:
Q6_K (GGUF)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagentic

User reports Ornith-1.5-9B at about 36 t/s generation and about 950 t/s prompt eval on an AMD RX 9060 XT 16GB with llama.cpp. Throughput drops to about 25 t/s generation and about 500 t/s prompt eval during long tasks. The run lasted about 3.5 hours on a coding agent task. The user contrasts this with Qwen 3.8 27B, which was frustratingly slow on the same hardware.

Sep 7, 2026

Get a weekly email of new AMD RX 9060 XT 16GB reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Qwen3.6 35B (3B active)
AMD RX 9060 XT 16GB
UD-Q4_K_XL
llama.cpp
40,96066.0 tokens/s
Qwen3.8 27B Swift-1.5
AMD RX 9060 XT 16GB
IQ3_S
Engine not reported
Not reported30.0 tokens/s
Qwen3.8 27B
AMD RX 9060 XT 16GB
IQ3_S
Engine not reported
Not reported50.0 tokens/s
Qwen2.5 14B Coder
AMD RX 9060 XT 16GB
Q4_K_M
llama.cpp
Not reported33.5 tokens/s
Ornith1.5 9B
AMD RX 9060 XT 16GB
Q6_K
llama.cpp
262,14436.0 tokens/s