llamaperf

AMD Radeon Pro W7900

AMD · 48GB · 1 report

Run models on your AMD Radeon Pro W7900? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the AMD Radeon Pro W7900

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 48 GB of VRAM.

This page is thin (1 of 3 reports needed for indexing). Help fill it in.
Tone: mixed
reported speed:
63.5 tokens/s generation · 1119.0 tokens/s prompt processing
quant:
UD-Q4_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingcreative-writing

User reports Qwen3.8-Flash-Next at 63.5 t/s decode (code) and 1119 t/s prefill at 2.6K context on a Radeon Pro W7900 eGPU plus Strix Halo iGPU hybrid setup. Setup is llama-halo-hybrid fork with UD-Q4_K_XL quant at 128K context, MTP head enabled, dense layers on the W7900 and routed experts on the iGPU. Hybrid beats gufo iGPU-only (44.7 t/s decode code, 903 t/s prefill at 2.6K) and leaves 43-52 GiB RAM free versus 14-30. GLM-5.3-Flash UD-IQ3_XXS does 25 t/s (31.6 with MTP) and DS-V4-Flash UD-IQ3_XXS does 22.5 t/s. W7900 runs 100-104 C junction on long prefills with a locked 241 W power cap. One run each, no quality testing.

Oct 6, 2026

Get a weekly email of new AMD Radeon Pro W7900 reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Qwen3.8 125B (6B active) Flash-Next
AMD Radeon Pro W7900
UD-Q4_K_XL
llama-halo-hybrid
128,00063.5 tokens/s