Qwen3.8 27B
on NVIDIA RTX 5090 · NInfer · 200,000 ctx
Sep 17, 2026
Summary
User reports Qwen3.8 27B at ~162.1 t/s generation and ~4.26k t/s prefill on an RTX 5090.
Setup is the Ninfer Windows Edition with a Q8 KV cache at 200k context.
Switching from LM Studio Q6 to Ninfer roughly doubled both decode and prompt speeds.