Qwen3.8 27B
on NVIDIA RTX 4090 · NInfer · 262,144 ctx
Oct 7, 2026
Use cases
coding
Summary
User reports Qwen3.8-27B at up to 149 tok/s code decode on one RTX 4090.
Setup is NInfer with INT8 prefill and E8 4-bit KV cache, full native 262K context, MTP3 speculative decoding, sm_89-retuned attention prefill, vision, and llama.cpp-compatible /metrics and /slots.
User says it is their go-to even with Strata, which is limited to IQ2_XS.