Qwen3.8 27B
on 4× NVIDIA RTX 3090 · SGLang · 8,192 ctx
Oct 3, 2026
Use cases
codingvisionlong-context
Summary
User reports Qwen3.8-27B at 136.8 tok/s single-stream decode (8k context, concurrency 1) on 4x RTX 3090.
Setup is SGLang with AWQ-INT4 (W4A16 Marlin) weights and bf16 KV cache, tensor parallel 4, DSpark speculative decoding enabled with a 1.36B draft model, 275,018-token KV pool.
Prefill sustains 1,581 tok/s at 8k c1 and 1.3-1.6k tok/s across all context lengths; aggregate decode reaches 169.8 tok/s at 8 concurrent requests. Max context 128k, max 9 concurrent requests (GDN state bound). Also benchmarks 2x 3090 (112.9 tok/s decode at 8k c1) and 1x 3090 without DSpark (43.2 tok/s decode, 8k context limit).