llamaperf

RTX Pro 4500 Blackwell 32GB

NVIDIA · 32GB · 1 report

See what fits on this GPU →

Use the calculator to check which models fit in 32 GB of VRAM. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →
This page is thin (1 of 3 reports needed for indexing). Help fill it in.
Tone: positive
reported speed:
45.2 tokens/s generation · 2022.5 tokens/s prompt processing
quant:
IQ4_XS (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks Qwen3.6 27B IQ4_XS on an RTX Pro 4500 Blackwell 32GB with llama.cpp b9007, reaching 2022.54 t/s prompt processing and 45.19 t/s generation. The same card also runs Qwen3.6 35B-A3B MXFP4 at 5507.10 t/s prompt processing and 159.81 t/s generation, along with Gemma4 26B-A4B MXFP4, Ernie 4.5 21B-A3B MXFP4, Nemotron Cascade 2 30B-A3B MXFP4, Tesselate OmniCoder 9B Q8, Qwen3.5 4B Q4_K, Qwen3.5 9B UD Q4_K_XL and GLM 4.7 Flash MXFP4. Compared with an RTX 5090, the 5090 is 60-70% faster at 2-3x power. User is happy with the card for 24/7 use.

Aug 28, 2026