Qwen3.6 27B
RTX Pro 4500 Blackwell 32GB · llama.cpp
- reported speed:
- 45.2 tokens/s generation · 2022.5 tokens/s prompt processing
- quant:
- IQ4_XS (gguf)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User benchmarks Qwen3.6 27B IQ4_XS on an RTX Pro 4500 Blackwell 32GB with llama.cpp b9007, reaching 2022.54 t/s prompt processing and 45.19 t/s generation. The same card also runs Qwen3.6 35B-A3B MXFP4 at 5507.10 t/s prompt processing and 159.81 t/s generation, along with Gemma4 26B-A4B MXFP4, Ernie 4.5 21B-A3B MXFP4, Nemotron Cascade 2 30B-A3B MXFP4, Tesselate OmniCoder 9B Q8, Qwen3.5 4B Q4_K, Qwen3.5 9B UD Q4_K_XL and GLM 4.7 Flash MXFP4. Compared with an RTX 5090, the 5090 is 60-70% faster at 2-3x power. User is happy with the card for 24/7 use.