llamaperf

AMD Radeon Pro W7800 48GB

AMD · 48GB · 1 report

As of 8 Oct 2026, the models most run on the AMD Radeon Pro W7800 48GB, with the median of plain runs (one device, one request, no speculative decoding):

Engines people use on it: llama.cpp 1

Run models on your AMD Radeon Pro W7800 48GB? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the AMD Radeon Pro W7800 48GB

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 48 GB of VRAM.

This page is thin (1 of 3 reports needed for indexing). Help fill it in.
Tone: mixed
reported speed:
90.0 tokens/s generation
quant:
IQ3_XXS (GGUF)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen 3.8 Flash-Next at about 90 t/s on a Radeon Pro W7800 48GB with 64GB DDR5 system RAM. Setup is llama.cpp with an IQ3_XXS GGUF quant, 128k context and Q8 KV cache, run with default settings. The user is satisfied with the speed but the model fell into repetition loops after 3-4 minutes of reasoning on a 27KB JavaScript audit task, repeating the same steps dozens of times; they interrupted it twice and ask whether a reasoning budget setting is the cause.

Oct 5, 2026

Get a weekly email of new AMD Radeon Pro W7800 48GB reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Qwen3.8 125B (6B active) Flash-Next
AMD Radeon Pro W7800 48GB
IQ3_XXS
llama.cpp
131,07290.0 tokens/s