llamaperf

Qwen3.8 27B

on NVIDIA RTX 5090 · NInfer

Tone: positive
Sep 18, 2026
Throughput
178.0 t/s gen
Quant
NVFP4

Summary

User reports Qwen3.8 27B NVFP4 at 178 t/s on an RTX 5090 after undervolting the GPU and capping CPU power at 50%. Setup uses the Ninfer inference engine. GPU memory was overclocked by 2400 MHz and the GPU was undervolted; CPU power was limited to 50% with no measurable TPS loss. A systemd service polls GPU usage and applies the CPU cap automatically during inference. Stock settings gave 172 t/s in synthetic tests; undervolting and VRAM overclocking raised this to 178 t/s, and DFlash2 speculative decoding currently yields 200 t/s. GPU power dropped from 600 W to under 450 W and GPU temperature from 75°C to 62°C.