llamaperf

Qwen3.8 27B QUASAR

on NVIDIA RTX 5090 · NInfer

Tone: positive
Oct 3, 2026
Throughput
287.0 t/s gen · 13700.0 t/s pp
Quant
NVFP4 (NVFP4)

Use cases

agenticcoding

Summary

User compares NInfer's official Qwen3.8-27B quant (part NVFP4, part FP8) against QUASAR's QAT full-NVFP4 checkpoint with DFlash2 embedded, both on a single RTX 5090. QUASAR reaches 484 tok/s on JSON output, 204 tok/s on prose, and 13.7k tok/s prefill, using 27.0 GB VRAM. The official quant gets 400 tok/s JSON, 176 tok/s prose, 11.4k tok/s prefill, and 30.8 GB VRAM. On a real agentic task the QUASAR build averages 287 tok/s over 33k tokens. GPQA Diamond scores 89.9% versus 90.4% for the official quant, and perplexity is about 1.9% worse.