llamaperf

Qwen3.8 27B

on 4× NVIDIA RTX 3090 · SGLang · 8,192 ctx

Oct 3, 2026
Throughput
136.8 t/s gen · 1581.0 t/s pp
Quant
AWQ-INT4 (AWQ)
KV cache
bf16
VRAM reported
24 GB

Use cases

codingvisionlong-context

Summary

User reports Qwen3.8-27B at 136.8 tok/s single-stream decode (8k context, concurrency 1) on 4x RTX 3090. Setup is SGLang with AWQ-INT4 (W4A16 Marlin) weights and bf16 KV cache, tensor parallel 4, DSpark speculative decoding enabled with a 1.36B draft model, 275,018-token KV pool. Prefill sustains 1,581 tok/s at 8k c1 and 1.3-1.6k tok/s across all context lengths; aggregate decode reaches 169.8 tok/s at 8 concurrent requests. Max context 128k, max 9 concurrent requests (GDN state bound). Also benchmarks 2x 3090 (112.9 tok/s decode at 8k c1) and 1x 3090 without DSpark (43.2 tok/s decode, 8k context limit).