llamaperf

Qwen3.8 27B

on NVIDIA RTX 5080 · NInfer · 131,072 ctx

Tone: positive
Sep 22, 2026
Throughput
71.6 t/s gen · 1380.6 t/s pp
Quant
Q3G64_F16S/Q4G64_F16S/Q5G64_F16S (ninfer)
KV cache
Q4 group64
VRAM reported
16 GB

Use cases

visionagenticlong-context

Summary

User reports Qwen3.8-27B at 71.57 t/s decode on a single RTX 5080 16GB at 131,072 context. Setup is NInfer v1.3 with mixed Q3/Q4/Q5 groupwise quantization (~3.95 BPW), Q4 group64 KV cache, MTP-3 speculative decoding, and Vision enabled. The 118,001-token prompt prefilled at 1380.61 t/s with 44.74% MTP acceptance. A short 512-token request averaged 84.7 t/s decode, and a video test reached 97.1 t/s decode with 100% MTP acceptance.