llamaperf

Radeon AI PRO R9700 32GB

AMD · 32GB · 16 reports

See what fits on this GPU →

Use the calculator to check which models fit in 32 GB of VRAM. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →
Tone: positive
reported speed:
61.4 tokens/s generation · 7800.0 tokens/s prompt processing
quant:
INT4 (W4A16)
kv:
FP8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Benchmark of Qwen3.6 35B-A3B MoE on single R9700 with vLLM. INT4 weights from Avesed. Also tested 27B dense with MTP spec=4. Prefill and decode speeds at various context depths. User is happy with results.

Tone: positive
reported speed:
120.0 tokens/s generation · 12000.0 tokens/s prompt processing
quant:
MXFP4-FP8
kv:
FP8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Post reports 80-120 t/s generation and 12k t/s prefill for single request on 4xR9700 with optimized vLLM image. Uses tcclaviger's MXFP4-FP8 quant. Command includes tensor-parallel-size 4, kv-cache-dtype fp8, speculative decoding with MTP.

quant:
Q4_K_XL (gguf)

User is asking for real-world sustained tok/s on R9700 with Qwen3.8-27B Q4 at 64K+ context. Mentions AMD blog quote of 51.8 tok/s and 5090 benchmarks showing degradation from 75 tok/s at 4K to ~26 tok/s at 64K. Also asks about MTP speculative decoding stability.

reported speed:
56.6 tokens/s generation · 198.5 tokens/s prompt processing
quant:
FP8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Two Radeon AI PRO R9700 GPUs (TP=2) on PCIe Gen5 x8, AMD Ryzen 5 7400 CPU, DDR5-6000 64GB. Running Qwen3.8-27B-FP8 with MTP3 speculative decoding. Generation throughput ~50-56 t/s, accepted ~30-37 t/s. User asks if numbers are reasonable and where to look for bottlenecks.

Tone: mixed
reported speed:
75.2 tokens/s generation · 671.9 tokens/s prompt processing
quant:
FP8
kv:
F16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Benchmark comparing PCIe gen4 x16+x4 vs gen5 x8/x8 for dual R9700 GPUs with tensor parallelism. Improvements of 18-32% in vLLM. Used stilldeadcode/vllm-radiance:0.7.4. CPU 9950X, 64GB DDR5 6000. Results shown for concurrency 1, 2, and 4.

Tone: mixed
reported speed:
40.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User compares Qwen3.8 27B to Qwen3.6 35B-A3B for daily coding use. Qwen3.8 is verbose and overthinks, while Qwen3.6 is faster and more concise. Mentions running on R9700 (Radeon AI PRO R9700) and hypothetical iGPU comparison.

reported speed:
26.6 tokens/s generation · 279.3 tokens/s prompt processing
quant:
Q4_K_XL (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Post describes using tensor-read-lazy with mmap to fit Q4 model on system with AMD Radeon AI PRO R9700, RX 9070 XT, and Radeon 8060S (Strix Halo). Model size 103.68 GiB, params 176.94 B. Prompt processing 279.25 t/s, generation 26.63 t/s.

Tone: positive
reported speed:
78.1 tokens/s generation · 1512.0 tokens/s prompt processing
quant:
IQ4_XS (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

visionlong-context

Custom kernels for RDNA4, ~3x generation speedup and ~30x prefill speedup over public vLLM-Radiance. Dual R9700s with 128GB RAM. Also mentions Muse Glimmer 30B dense model work.

Qwen3.8 27B

Radeon AI PRO R9700 32GB · llama.cpp · 255,000 ctx

Tone: positive
reported speed:
51.6 tokens/s generation
quant:
Q8_K_XL (gguf)
kv:
F16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Running Qwen3.8 27B Q8_K_XL with MTP on 2x Radeon AI PRO R9700 (32GB each) using llama.cpp with ROCm 10.0. Reported 37-50 tg/s, spiking over 60 tg/s when writing code. Context length 255k, KV cache F16. Draft acceptance 0.72.

Qwen3.8 27B

Radeon AI PRO R9700 32GB · vLLM · 94,065 ctx

Tone: positive
reported speed:
280.0 tokens/s generation · 3831.0 tokens/s prompt processing
quant:
MXFP4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

The post reports 280 tok/s decode on Qwen3.8 27B with MXFP4 quantization on dual R9700s. The prompt processing speed at 128k context is 3831 t/s. The author expresses enthusiasm about the performance and collaboration.

Qwen3.8 27B

Radeon AI PRO R9700 32GB · vLLM-Radiance · 262,000 ctx

Tone: positive
reported speed:
36.6 tokens/s generation · 17636.0 tokens/s prompt processing
quant:
FP8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Quad R9700 setup with vLLM-Radiance. Prefill 17636 t/s, TG 36.6 t/s (best TG 106 t/s with 80% MTP 4 acceptance). Running Hermes with Qwen 3.8 27b fp8 262k ctx. One GPU on PCIe 4x8, three on 4x16. vLLM-Radiance works well with quad setup despite official dual support.

reported speed:
29.0 tokens/s generation
quant:
Q4_K_XL (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Theoretical max TG/s calculation for dense models. Example: Qwen3.8 27B Q4_K_XL on Radeon AI PRO R9700 with llama.cpp yields 29 TG/s (76% of theoretical 38 TG/s). Also mentions RTX 5090 with 1.8 TB/s bandwidth.

Qwen3.8 27B

Radeon AI PRO R9700 32GB · vLLM · 16,384 ctx

Tone: positive
reported speed:
87.6 tokens/s generation · 4134.0 tokens/s prompt processing
quant:
FP8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Benchmark of Qwen3.8-27B FP8 on 2x R9700 with vLLM Radiance TP2. Also tested MXFP4 (111.4 t/s) and Qwen3.8 Flash Next (35.4 t/s).

reported speed:
84.5 tokens/s generation · 2000.4 tokens/s prompt processing
quant:
Q5_K_M (gguf)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingtool-useagentic

Qwen3.6-35b is my daily driver. Reliable and fast. I've tried ROCM and the PP was much faster, but gen was slower. I'm testing MTP now, and the PP is slower at 1700 t/s but gen is around 110 - 130 t/s. command: > --model /models/Qwen3.6-35B-A3B-UD-Q5_K_M.gguf -ngl 99 -lv 4 --split-mode layer --ctx-size 262144 --cache-type-k q8_0 --cache-type-v q8_0 --threads 8 --metrics --parallel 1 --batch_size 2048 --ubatch_size 2048 --no-mmap --cache-ram 0 --temp 0.6 --top-p 0.95 --top-k 40 --min-p 0.08 -fa on --chat-template-kwargs '{"preserve_thinking":false}'

Qwen3.6 27B

Radeon AI PRO R9700 32GB · llama.cpp · 131,072 ctx

quant:
Q8_0 (gguf)
kv:
F16
flash attention:
on
mtp (multi-token prediction):
on
codingsummarizationlong-context

Multi-GPU setup with 2x Radeon AI PRO R9700 32GB. Decode t/s ranged from 40-67 depending on context length. Prefill throughput 410-1500 t/s. MTP draft acceptance 0.33-0.61. KV cache in F16. Context 131072.

Tone: positive
reported speed:
51.1 tokens/s generation
quant:
q2

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Lucebox custom setup with AMD Radeon AI PRO R9700 and Strix Halo 128GB. Asymmetric parallelism: R9700 handles dense path, hot experts, cache, draft model; Strix Halo holds other experts. 3.63x faster than single DGX Spark. Experimental, q2 quant, 16k context. Working on KVFlash for 64k-128k.