Qwen3.8 27B
on NVIDIA RTX 5080 · NInfer · 110,592 ctx
Oct 6, 2026
Use cases
codingtool-uselong-context
Summary
User reports Qwen 3.8 27B at a median 96.45 t/s decode on a 16GB RTX 5080, with a range of roughly 90-110 t/s during coding tasks.
Setup is NInfer v1.5 with q4 KV cache, MTP speculative decoding with 3 draft tokens, 110,592 context allocated, and 11.86 GiB of weights leaving about 498 MiB of slack.
Across 32 completed requests the lowest decode was 84.6 t/s and the highest 131.7 t/s, with median time to first token 1.4 seconds. About 68% of generated tokens were reasoning tokens, and one request that lost its prompt cache took almost 50 seconds to process a 79k prompt before generating at about 94 t/s.