llamaperf

Qwen3.8 27B

on NVIDIA RTX Pro 6000 Blackwell · SGLang · 262,144 ctx

Tone: positive
Sep 26, 2026
Throughput
77.6 t/s gen · 14-37 t/s pp
Quant
BF16 (safetensors)
KV cache
FP8
VRAM reported
96 GB

Use cases

agentictool-uselong-context

Summary

User reports Qwen3.8-27B at 77-80 tok/s on a single RTX PRO 6000 Blackwell 96GB. Setup is SGLang with BF16 safetensors, FP8 KV cache, EAGLE speculative decoding, and 262144 context length. The decode figure is a range; EAGLE accept rate was 0.94-1.00 and the user claims about 2.2x speedup over non-speculative decoding. Prefill was 14-37 tok/s on 122-token batches.