Qwen3.8 27B
on 2× NVIDIA RTX 2080 Ti 22GB (modded) · vLLM · 524,288 ctx
Oct 7, 2026
Use cases
codingagentictool-usevisionlong-context
Summary
User reports Qwen3.8-27B at ~59-68 tok/s single-stream decode on 2x modded RTX 2080 Ti 22GB (SM75, TP=2, NVLink).
Setup is a vLLM fork with GPTQ-Int4 self-quant, turboquant_k3v4_nc KV cache, 524,288 max-model-len (YaRN 2x), MTP K=2 speculative decoding, and vision enabled.
Aggregate throughput is 364.7 tok/s at 16 lanes; 24 lanes regresses. The S4 scoped re-emission drafter measured 811.5 tok/s on copy-shaped spans but crashes at current HEAD, and v3 GPU-merge async is negative.