Qwen3.8 27B
on NVIDIA RTX 4090 · NInfer · 262,144 ctx
Sep 23, 2026
Use cases
coding
Summary
User reports Qwen3.8-27B at 148.6 tok/s on a single RTX 4090 with MTP3 speculative decoding at 81.0% draft acceptance.
Setup is NInfer with the official groupwise artifact, INT8 KV cache, CUDA Graphs on, prefill-chunk 1024, single request greedy decoding.
Without speculation decode is 50.5 tok/s; at 128K context depth it is 39.6 tok/s. Prefill is 1,849 tok/s at 64K and 1,561 tok/s at 128K. The E8 4-bit KV default profile costs about 5.7% decode versus INT8, giving roughly 126.6 tok/s on the code probe.