Qwen3.8 27B
on NVIDIA RTX 5090 · NInfer · 240,000 ctx
Sep 27, 2026
Use cases
long-contextsummarizationtool-useagentic
Summary
User reports Qwen3.8-27B at 158 tok/s decode and 7,265 tok/s prefill on an RTX 5090 32GB (eGPU via OCuLink Gen4 x4).
Setup is NInfer with NVFP4 weights, FP8 KV cache, 240K context, MTP3 speculative decoding at 76% acceptance, and 2 concurrent lanes. Decode rises to 213 tok/s at 32K and 202 tok/s at 128K; prefill is 6,892 tok/s at 32K and 3,904 tok/s at 128K.
User compares against llama.cpp (Q5_K_M GGUF, q8_0 KV, 196K context, MTP on) at 114 tok/s decode and 1,545 tok/s prefill at 1K, and notes vLLM at about 70 tok/s decode at short context. Quality was statistically indistinguishable across engines on a 250+ item eval. NInfer does not support json_mode.