llamaperf

Qwen3.8 27B

on NVIDIA RTX 5090 · NInfer · 200,000 ctx

Tone: positive
Sep 17, 2026
Throughput
162.1 t/s gen · 4260.0 t/s pp
Quant
Q8
KV cache
Q8

Summary

User reports Qwen3.8 27B at ~162.1 t/s generation and ~4.26k t/s prefill on an RTX 5090. Setup is the Ninfer Windows Edition with a Q8 KV cache at 200k context. Switching from LM Studio Q6 to Ninfer roughly doubled both decode and prompt speeds.