llamaperf

Qwen3.8 27B

on NVIDIA RTX 4090 · NInfer · 262,144 ctx

Sep 23, 2026
Throughput
148.6 t/s gen · 1849.0 t/s pp
Quant
E8 4-bit
KV cache
INT8
VRAM reported
24 GB

Use cases

coding

Summary

User reports Qwen3.8-27B at 148.6 tok/s on a single RTX 4090 with MTP3 speculative decoding at 81.0% draft acceptance. Setup is NInfer with the official groupwise artifact, INT8 KV cache, CUDA Graphs on, prefill-chunk 1024, single request greedy decoding. Without speculation decode is 50.5 tok/s; at 128K context depth it is 39.6 tok/s. Prefill is 1,849 tok/s at 64K and 1,561 tok/s at 128K. The E8 4-bit KV default profile costs about 5.7% decode versus INT8, giving roughly 126.6 tok/s on the code probe.