llamaperf

AMD RX 7600 XT 16GB

AMD · 16GB · 3 reports

As of 7 Oct 2026, the models most run on the AMD RX 7600 XT 16GB, with the median of plain runs (one device, one request, no speculative decoding):

Engines people use on it: llama.cpp 2

Run models on your AMD RX 7600 XT 16GB? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the AMD RX 7600 XT 16GB

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 16 GB of VRAM.

Qwen3.8 27B

AMD RX 7600 XT 16GB · llama.cpp · 114,688 ctx

Tone: positive
quant:
GSQ-RCO-IQ3_XXS (GGUF)
kv:
Q4_0
mtp (multi-token prediction):
on
codingagentic

User benchmarks Qwen3.8 27B on an RX 7600 XT 16GB, completing 15/15 tasks with 1.000 correctness in 348.5 seconds. Setup is llama.cpp HIP ROCm with the GSQ-RCO-IQ3_XXS GGUF and a Q4_0 KV cache at 114,688 context, fully in VRAM with no CPU offload. The same benchmark also ran Ornith-1.5-9B (15/15, 0.983, 184.9s) and K2-Horizon-7B (11/15, 0.909, 498.1s); the user calls Qwen3.8 27B the undisputed winner for agentic coding.

Sep 23, 2026

Qwen3.8 27B

AMD RX 7600 XT 16GB · llama.cpp · 163,840 ctx

Tone: positive
reported speed:
18.0 tokens/s generation · 141.0 tokens/s prompt processing
quant:
IQ3_XXS (GGUF)
kv:
q8_0
mtp (multi-token prediction):
off

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen3.8 27B at 18 t/s decode and 141 t/s prefill on a 1.8k-token prompt, running on an AMD Radeon RX 7600 XT 16 GB at 163,840 context. Setup is llama.cpp llama-server build 10480 with the unsloth Qwen3.8-27B-UD-IQ3_XXS GGUF (10.2 GiB), q8_0 KV cache, flash attention on, and Vulkan (RADV) backend. The model is a hybrid architecture where only 16 of 64 layers keep a full KV cache, so KV is about 5.3 GiB at this context. With MTP on the same model runs about 24 t/s at 98k context and 39 t/s at 124k. The user notes 160k is a VRAM-math ceiling, not a quality claim, and that decode drops to about 6 t/s if the GPU spills to system RAM.

Sep 23, 2026

Muse 30B Glimmer

AMD RX 7600 XT 16GB · llama.cpp · 62,144 ctx

Tone: positive
reported speed:
20.0 tokens/s generation · 308.0 tokens/s prompt processing
quant:
UD-Q2-K-XL

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Muse Glimmer 30B at about 308 t/s prompt processing and about 20 t/s generation on an RX 7600 XT 16GB. Setup is llama.cpp with ROCm, the UD-Q2-K-XL quant and DFlash speculative decoding. The run completed a coding task.

Sep 8, 2026

Get a weekly email of new AMD RX 7600 XT 16GB reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Qwen3.8 27B
AMD RX 7600 XT 16GB
IQ3_XXS
llama.cpp
163,84018.0 tokens/s
Muse 30B Glimmer
AMD RX 7600 XT 16GB
UD-Q2-K-XL
llama.cpp
62,14420.0 tokens/s