Qwen3.8 27B
on NVIDIA RTX 4090 · SGLang
Oct 3, 2026
Summary
User reports Qwen3.8-27B at 135.4 tok/s decode on a bare RTX 4090 24 GB.
Setup is SGLang with EXL3 3bpw weights, 207,356-token KV pool, 22.9 GB peak, 239 ms TTFT over 3 samples.
A separate depth sweep on the same card gives 184 tok/s code and 126 tok/s prose at ~50 prompt tokens, falling to 115 and 91 tok/s at 133k-190k tokens; cold prefill was 2,112 tok/s at 19k and 1,636 at 134k. The recipe is the unchanged RTX 3090 SGLang EXL3 config, and it now displaces a TabbyAPI SC4bpw recipe that measured 86 tok/s. No coherence ladder, soak test or 3 bpw vs 4 bpw quality comparison was done.