Qwen2.5 32B
2× NVIDIA RTX Pro 4000 Blackwell · SGLang
- reported speed:
- 71.4 tokens/s generation
- quant:
- AWQ (AWQ)
- kv:
- fp8_e5m2
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User benchmarks Qwen2.5-32B-Instruct-AWQ at 71.37 tok/s median decode on 2x RTX PRO 4000 Blackwell. Setup is SGLang 0.5.19 with AWQ weights, NGRAM speculative decoding (K=5), tensor parallelism 2 over PCIe Gen4, and FP8 KV cache. NGRAM speculative decoding gives a 1.28x median speedup over the 55.60 tok/s baseline, with high inter-prompt variance (std dev 19.18 tok/s). Per-domain decode ranges are 65-107 tok/s for JSON, 60-226 tok/s for code, and 58-82 tok/s for prose. FP8 KV cache alone measured 54.77 tok/s, a 1.5% regression.