llamaperf

Qwen2.5 32B

on 2× NVIDIA RTX Pro 4000 Blackwell · SGLang

Tone: positive
Sep 26, 2026
Throughput
71.4 t/s gen
Quant
AWQ (AWQ)
KV cache
fp8_e5m2
VRAM reported
24 GB

Use cases

codinglong-context

Summary

User benchmarks Qwen2.5-32B-Instruct-AWQ at 71.37 tok/s median decode on 2x RTX PRO 4000 Blackwell. Setup is SGLang 0.5.19 with AWQ weights, NGRAM speculative decoding (K=5), tensor parallelism 2 over PCIe Gen4, and FP8 KV cache. NGRAM speculative decoding gives a 1.28x median speedup over the 55.60 tok/s baseline, with high inter-prompt variance (std dev 19.18 tok/s). Per-domain decode ranges are 65-107 tok/s for JSON, 60-226 tok/s for code, and 58-82 tok/s for prose. FP8 KV cache alone measured 54.77 tok/s, a 1.5% regression.