llamaperf

Qwen3.8 27B

on NVIDIA RTX 5090 · NInfer · 240,000 ctx

Tone: positive
Sep 27, 2026
Throughput
158.0 t/s gen · 7265.0 t/s pp
Quant
NVFP4 (NVFP4)
KV cache
FP8
System RAM
32 GB
VRAM reported
32 GB

Use cases

long-contextsummarizationtool-useagentic

Summary

User reports Qwen3.8-27B at 158 tok/s decode and 7,265 tok/s prefill on an RTX 5090 32GB (eGPU via OCuLink Gen4 x4). Setup is NInfer with NVFP4 weights, FP8 KV cache, 240K context, MTP3 speculative decoding at 76% acceptance, and 2 concurrent lanes. Decode rises to 213 tok/s at 32K and 202 tok/s at 128K; prefill is 6,892 tok/s at 32K and 3,904 tok/s at 128K. User compares against llama.cpp (Q5_K_M GGUF, q8_0 KV, 196K context, MTP on) at 114 tok/s decode and 1,545 tok/s prefill at 1K, and notes vLLM at about 70 tok/s decode at short context. Quality was statistically indistinguishable across engines on a 250+ item eval. NInfer does not support json_mode.