llamaperf

Qwen3.8 27B Uncensored

on NVIDIA RTX 5090 · NInfer · 262,144 ctx

Tone: positive
Sep 24, 2026
Throughput
175.0 t/s gen
Quant
NVFP4
KV cache
Q4
MTP (Multi-Token Prediction)
on
VRAM reported
32 GB

Use cases

agenticcodinglong-context

Summary

User reports Qwen3.8 27B uncensored at 175 t/s on an RTX 5090 32GB. Setup is NInfer with NVFP4 / groupwise-int quantization and MTP enabled, running at 262,144 context with a Q4 KV cache. The user says the Q4 KV cache matched Q8 quality and that the longer context enabled long reasoning tasks in a Hermes harness.