Qwen3.8 27B
on NVIDIA RTX 3090 · SGLang · 131,072 ctx
Oct 3, 2026
Summary
User reports Qwen3.8-27B at 98.2 t/s on prose and 225.1 t/s on code on one RTX 3090.
Setup is SGLang with the sglang-exl3 plugin, EXL3 3.00bpw weights, fp8 KV cache, a 131,072-token window, and DFlash2 speculative decoding with a 5.0bpw draft.
Prompt processing on a new uncached 32k prompt runs at about 914 tok/s with 35.9 s to first token. Compared with the MTP recipe on the same card, code is +57% (225 vs 143) and prose is the same (98 vs 98), with a 131k window instead of 262k. Max concurrency is 1 because two 32k streams stall each other's decode.