llamaperf

Qwen3.8 27B

on NVIDIA RTX 3090 · NInfer · 8,192 ctx

Oct 2, 2026
Throughput
71.0 t/s gen · 861.5 t/s pp
KV cache
INT8
VRAM reported
24 GB

Summary

User reports Qwen3.8-27B at 71.00 tok/s decode on one RTX 3090 24 GB. Setup is NInfer v0.6.1 with INT8 KV cache, ReplaySSM and MTP3 speculative decoding, CUDA Graphs, 8,192-token context, 19.6 GB VRAM used. Decode by cohort: C2 90.66, C4 100.28, C8 165.33 aggregate tok/s. Prefill with 4,362 fresh input tokens per request: C1 861.51 tok/s (TTFT 4,893 ms), C8 844.10 tok/s. MTP3 acceptance 61.13%, C1 TTFT 149 ms.