llamaperf

AMD Radeon 890M vs AMD Radeon AI PRO R9700 32GB for local LLMs

1 report on the AMD Radeon 890M and 50 on the AMD Radeon AI PRO R9700 32GB

Which is faster for local LLMs?

On Qwen3.5 35B · 3B active at 4-bit, the AMD Radeon 890M runs at 14.9 tokens per second (one run) and the AMD Radeon AI PRO R9700 32GB at 149.2 (one run). Those are community reports whose engines and settings differ, so the gap is a rough sign, not a controlled measure of the two cards.

AMD Radeon 890M
Memory
shared memory unified
Memory bandwidth
not on record
FP16 compute
not on record
Reports
1
AMD Radeon AI PRO R9700 32GB
Memory
32GB
Memory bandwidth
640 GB/s
FP16 compute
191 TFLOPS
Reports
50

Measured on both

Median tokens per second from plain runs (one device, one request, no speculative decoding, the whole model in the device's memory) of the same model size at the same quant level. Engines and context lengths can still differ between the runs. How to read these

ModelAMD Radeon 890MAMD Radeon AI PRO R9700 32GB
Qwen3.5 35B · 3B active4-bit14.9one run149.2one run

Estimated by the calculator

The quant the calculator recommends for each card and its estimated speed at a 32,768-token context with 32 GB of system RAM, from memory bandwidth and the model's shape. Where a model was also measured above, the measurement wins.

ModelAMD Radeon 890MAMD Radeon AI PRO R9700 32GB
Qwen3.5 35B · 3B activenot in the calculator yet58.2 t/sQ5_K_M
Qwen3.8 27Bnot in the calculator yet16.0 t/sQ6_K
Qwen3.8 125B · 6B activenot in the calculator yet18.7 t/sQ2_K, experts in RAM
DeepSeek V4 Flash 284B · 13B activenot in the calculator yetdoesn't fit at these settings
Qwen3.6 35B · 3B activenot in the calculator yet58.2 t/sQ5_K_M
Qwen3.6 27Bnot in the calculator yet16.0 t/sQ6_K
Gemma 4 26B · 4B activenot in the calculator yet40.9 t/sQ6_K

Change the context, RAM or model in the calculator: AMD Radeon 890M · AMD Radeon AI PRO R9700 32GB

Every report on each card, with its source: AMD Radeon 890M · AMD Radeon AI PRO R9700 32GB

More comparisons