llamaperf

AMD Radeon 890M

AMD · shared memory · 1 report

As of 9 Oct 2026, the models most run on the AMD Radeon 890M, with the median of plain runs (one device, one request, no speculative decoding):

Engines people use on it: llama.cpp 1

Run models on your AMD Radeon 890M? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the AMD Radeon 890M

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. It has no memory of its own and borrows the machine's RAM, so how large a model fits depends on the machine, and the calculator does not size it.

This page is thin (1 of 3 reports needed for indexing). Help fill it in.
reported speed:
14.8 tokens/s generation · 212.1 tokens/s prompt processing
quant:
Q4_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks Qwen3.5-35B-A3B at 14.85 t/s generation and 212.05 t/s prompt processing on an AMD Radeon 890M (gfx1150) iGPU. Setup is llama.cpp (build a0ed91a44) with a Q4_K_XL GGUF, 99 layers offloaded, comparing Vulkan and ROCm backends with flash attention on and off. ROCm reaches 227.54 t/s prefill and 12.68 t/s decode without flash attention, and 229.81 t/s prefill with 13.18 t/s decode with it. A second build with unified memory and ROCm flash attention gave similar results. The same machine also ran llama-2-7b Q4_0 at 17.14 t/s decode on Vulkan and 15.52 t/s on ROCm.

Oct 9, 2026

Get a weekly email of new AMD Radeon 890M reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Qwen3.5 35B (3B active)
AMD Radeon 890M
Q4_K_XL
llama.cpp
Not reported14.8 tokens/s