llamaperf

Qwen3.8 27B

on NVIDIA RTX 4090 · NInfer · 262,144 ctx

Tone: positive
Oct 7, 2026
Throughput
149.0 t/s gen
Quant
INT8
KV cache
E8 4-bit

Use cases

coding

Summary

User reports Qwen3.8-27B at up to 149 tok/s code decode on one RTX 4090. Setup is NInfer with INT8 prefill and E8 4-bit KV cache, full native 262K context, MTP3 speculative decoding, sm_89-retuned attention prefill, vision, and llama.cpp-compatible /metrics and /slots. User says it is their go-to even with Strata, which is limited to IQ2_XS.