Qwen3.8 27B Uncensored
on NVIDIA RTX 5090 · NInfer · 262,144 ctx
Sep 24, 2026
Use cases
agenticcodinglong-context
Summary
User reports Qwen3.8 27B uncensored at 175 t/s on an RTX 5090 32GB.
Setup is NInfer with NVFP4 / groupwise-int quantization and MTP enabled, running at 262,144 context with a Q4 KV cache.
The user says the Q4 KV cache matched Q8 quality and that the longer context enabled long reasoning tasks in a Hermes harness.