Qwen3.8 27B QUASAR
on NVIDIA RTX 5090 · NInfer
Oct 3, 2026
Use cases
agenticcoding
Summary
User compares NInfer's official Qwen3.8-27B quant (part NVFP4, part FP8) against QUASAR's QAT full-NVFP4 checkpoint with DFlash2 embedded, both on a single RTX 5090.
QUASAR reaches 484 tok/s on JSON output, 204 tok/s on prose, and 13.7k tok/s prefill, using 27.0 GB VRAM. The official quant gets 400 tok/s JSON, 176 tok/s prose, 11.4k tok/s prefill, and 30.8 GB VRAM.
On a real agentic task the QUASAR build averages 287 tok/s over 33k tokens. GPQA Diamond scores 89.9% versus 90.4% for the official quant, and perplexity is about 1.9% worse.