llamaperf

Qwen3.8 27B

on NVIDIA RTX 4090 · NInfer · 262,000 ctx

Tone: positive
Sep 27, 2026
Throughput
107.0 t/s gen · 5008.0 t/s pp
Quant
Q4/Q5
KV cache
rk4v4-e8

Use cases

codingtool-uselong-contextvision

Summary

User reports Qwen3.8-27B at 107 tok/s decode with MTP on mixed prompts and 5,008 tok/s prefill at 2K tokens on one RTX 4090 under native Windows. Setup is NInfer with Q4/Q5 weights, rk4v4-e8 KV cache, 262K context per request, MTP with 3 draft tokens and n-gram chaining. Decode reaches 289 tok/s with MTP + n-gram on edit-style prompts and 211 tok/s with DFlash2 on code; prose with thinking on gives 99 tok/s with DFlash2 versus 85 tok/s with MTP + n-gram. Against llama.cpp on the same card, prefill is 5,008 vs 2,729 tok/s at 2K, a 64K prompt takes 17.2 s vs 29.8 s and a 128K prompt 43.0 s vs 73.6 s. Ternary Bonsai 2 27B reaches 188 tok/s decode on mixed prompts, 532 tok/s with MTP + n-gram on edits and 360 tok/s aggregate with three concurrent requests. The 4090 also drives a 4K monitor at 60 Hz, costing about 15% of decode.