Qwen3.8 27B QUASAR
NVIDIA RTX 5090 · NInfer
- reported speed:
- 287.0 tokens/s generation · 13700.0 tokens/s prompt processing
- quant:
- NVFP4 (NVFP4)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User compares NInfer's official Qwen3.8-27B quant (part NVFP4, part FP8) against QUASAR's QAT full-NVFP4 checkpoint with DFlash2 embedded, both on a single RTX 5090. QUASAR reaches 484 tok/s on JSON output, 204 tok/s on prose, and 13.7k tok/s prefill, using 27.0 GB VRAM. The official quant gets 400 tok/s JSON, 176 tok/s prose, 11.4k tok/s prefill, and 30.8 GB VRAM. On a real agentic task the QUASAR build averages 287 tok/s over 33k tokens. GPQA Diamond scores 89.9% versus 90.4% for the official quant, and perplexity is about 1.9% worse.