llamaperf

Qwen3.8 27B

on 2× NVIDIA RTX 2080 Ti 22GB (modded) · vLLM · 524,288 ctx

Tone: positive
Oct 7, 2026
Throughput
59-68 t/s gen
Quant
GPTQ-Int4 (GPTQ)
KV cache
turboquant_k3v4_nc
VRAM reported
22 GB

Use cases

codingagentictool-usevisionlong-context

Summary

User reports Qwen3.8-27B at ~59-68 tok/s single-stream decode on 2x modded RTX 2080 Ti 22GB (SM75, TP=2, NVLink). Setup is a vLLM fork with GPTQ-Int4 self-quant, turboquant_k3v4_nc KV cache, 524,288 max-model-len (YaRN 2x), MTP K=2 speculative decoding, and vision enabled. Aggregate throughput is 364.7 tok/s at 16 lanes; 24 lanes regresses. The S4 scoped re-emission drafter measured 811.5 tok/s on copy-shaped spans but crashes at current HEAD, and v3 GPU-merge async is negative.