llamaperf

Qwen3.8 27B

on NVIDIA RTX Pro 6000 Blackwell · SGLang · 262,144 ctx

Tone: positive
Oct 3, 2026
Throughput
210.0 t/s gen
Quant
NVFP4 (NVFP4)
System RAM
96 GB
VRAM reported
96 GB

Use cases

agentictool-uselong-contextcodingvision

Summary

User benchmarks Qwen3.8-27B NVFP4 at 210 tok/s single-stream on one RTX PRO 6000 Blackwell, using SGLang with the DFlash2 drafter. Setup is SGLang v0.5.20 with a 262,144-token window and 4 slots; the same model without a drafter ran 75 tok/s, and vLLM with DFlash2 ran 160 tok/s. The 27B matrix ran at a 400W power cap. User also reports 607 tok/s aggregate across 4 users, 97s full-window prefill, 73.3% BFCL core, 700/1000 arena score, and 24/40 CAPTCHA puzzles. Flash-Next and an uncensored 27B fine-tune were tested alongside.