llamaperf

Qwen3.8 27B

on NVIDIA CMP 170HX 64GB (unlocked) · vLLM · 250,000 ctx

Tone: positive
Sep 20, 2026
Throughput
147.0 t/s gen
Quant
W4A16 (W4A16)
KV cache
FP8
VRAM reported
64 GB

Use cases

long-context

Summary

User reports Qwen3.8-27B at 147.0 t/s single-stream decode on one NVIDIA CMP 170HX 64GB, with 134.7 t/s at 4K, 100.1 t/s at 65K, ~90 t/s at 126K, and 64.9 t/s at 250K context. Setup is vLLM 0.27.1 with a W4A16 target, a DFlash2 W4A16 drafter, FP8 target KV, BF16 draft KV, a custom SM80 split-KV verifier, full CUDA Graph, 35 verifier segments / 140 CTAs, k=3 draft tokens, MAX_SEQS=1, and 1350 MHz / 180W. These are decode-only numbers, not end-to-end throughput including prefill. FP8 beat INT8 at long context (48.3 vs 115.4 ms/iter at 250K). Allowing 2 concurrent long requests made makespan and slowest-request throughput worse at both 126K and 250K.