llamaperf

AMD RX 6900 XT 16GB

AMD · 16GB · 2 reports

Engines people use on it: Strata 1

Run models on your AMD RX 6900 XT 16GB? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the AMD RX 6900 XT 16GB

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 16 GB of VRAM.

This page is thin (2 of 3 reports needed for indexing). Help fill it in.
Tone: mixed
reported speed:
25.0 tokens/s generation
quant:
GSQ-RCO

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8-Flash-Next at 25 tok/s on an RX 6900 XT installed in a retired HP DL380p server with 172 GB DDR3 and 2x 10-core CPUs. Setup is Strata with llama.cpp and an HTTP router, GSQ-RCO quant at 64k context. The GPU is powered by an external desktop PSU; total draw is 250 W idle and about 450 W while inferring. User also reports Qwen3.8-27B with GSQ-RCO-IQ3_XXS and MTP at 192k context reaching 45 tok/s. The machine is a prototype; user plans a flexible riser and a proper GPU platform, and may add Nvidia P40s in the remaining slots someday.

Oct 6, 2026
Tone: positive
reported speed:
30-45 tokens/s generation
quant:
IQ3_XXS (GGUF)
kv:
q8_0/q4_0
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports running Qwen3.8-27b (Huihui abliterated, IQ3_XXS) on an RX 6900 XT 16GB with llama.cpp and ROCm 10. Prefill drops from 400 t/s to 260 t/s at 40k context; decode varies between 30 t/s and 45 t/s. Setup uses 100k context, KV cache q8_0/q4_0, flash attention, ngram-mod, MTP with n=2, and mmproj in system RAM; VRAM fills to 15.7/16 GB.

Sep 26, 2026

Get a weekly email of new AMD RX 6900 XT 16GB reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Qwen3.8 125B (6B active) Flash-Next
AMD RX 6900 XT 16GB
GSQ-RCO
Strata
64,00025.0 tokens/s