Qwen3.8 27B
on NVIDIA RTX 5080 · NInfer · 131,072 ctx
Sep 22, 2026
Use cases
visionagenticlong-context
Summary
User reports Qwen3.8-27B at 71.57 t/s decode on a single RTX 5080 16GB at 131,072 context.
Setup is NInfer v1.3 with mixed Q3/Q4/Q5 groupwise quantization (~3.95 BPW), Q4 group64 KV cache, MTP-3 speculative decoding, and Vision enabled. The 118,001-token prompt prefilled at 1380.61 t/s with 44.74% MTP acceptance.
A short 512-token request averaged 84.7 t/s decode, and a video test reached 97.1 t/s decode with 100% MTP acceptance.