llamaperf

NVIDIA RTX Pro 4000 Blackwell

NVIDIA · 24GB · 3 reports

Engines people use on it: NInfer 1 · Ollama 1 · SGLang 1

Run models on your NVIDIA RTX Pro 4000 Blackwell? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the NVIDIA RTX Pro 4000 Blackwell

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 24 GB of VRAM.

Tone: positive
reported speed:
71.4 tokens/s generation
quant:
AWQ (AWQ)
kv:
fp8_e5m2

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codinglong-context

User benchmarks Qwen2.5-32B-Instruct-AWQ at 71.37 tok/s median decode on 2x RTX PRO 4000 Blackwell. Setup is SGLang 0.5.19 with AWQ weights, NGRAM speculative decoding (K=5), tensor parallelism 2 over PCIe Gen4, and FP8 KV cache. NGRAM speculative decoding gives a 1.28x median speedup over the 55.60 tok/s baseline, with high inter-prompt variance (std dev 19.18 tok/s). Per-domain decode ranges are 65-107 tok/s for JSON, 60-226 tok/s for code, and 58-82 tok/s for prose. FP8 KV cache alone measured 54.77 tok/s, a 1.5% regression.

Sep 26, 2026
Tone: positive
reported speed:
67.0 tokens/s generation · 785.0 tokens/s prompt processing
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports 67 t/s with MTP3 speculative decoding enabled, against 24.4 tok/s without MTP. Setup uses an INT8 KV cache with group-64. The figure comes from a 128K NIAH benchmark with 130,048 prompt tokens, where MTP acceptance was 100% on a deterministic answer. Roughly 727 MiB of VRAM was left.

Sep 7, 2026
Tone: positive
reported speed:
33.9 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8:27b at 33.91 t/s on 2x RTX PRO 4000s and an RTX PRO 2000, up from 12.2 t/s. Splitting the GPUs across VMs produced the gain. User also reports muse-glimmer:30B at 22.61 t/s on the same setup, up from 14.3 t/s.

Sep 7, 2026

Get a weekly email of new NVIDIA RTX Pro 4000 Blackwell reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Qwen2.5 32B
2× NVIDIA RTX Pro 4000 Blackwell
AWQ
SGLang
Not reported71.4 tokens/s
Qwen3.8 27B
NVIDIA RTX Pro 4000 Blackwell
Not reported
NInfer
128,00067.0 tokens/s
Qwen3.8 27B
NVIDIA RTX Pro 4000 Blackwell
Not reported
Ollama
Not reported33.9 tokens/s