Qwen2.5 32B
on 2× NVIDIA RTX Pro 4000 Blackwell · SGLang
Sep 26, 2026
Use cases
codinglong-context
Summary
User benchmarks Qwen2.5-32B-Instruct-AWQ at 71.37 tok/s median decode on 2x RTX PRO 4000 Blackwell.
Setup is SGLang 0.5.19 with AWQ weights, NGRAM speculative decoding (K=5), tensor parallelism 2 over PCIe Gen4, and FP8 KV cache.
NGRAM speculative decoding gives a 1.28x median speedup over the 55.60 tok/s baseline, with high inter-prompt variance (std dev 19.18 tok/s). Per-domain decode ranges are 65-107 tok/s for JSON, 60-226 tok/s for code, and 58-82 tok/s for prose. FP8 KV cache alone measured 54.77 tok/s, a 1.5% regression.