llamaperf

Qwen3.6 27B

on 2× NVIDIA RTX 2080 Ti 22GB (modded) · vLLM · 4,096 ctx

Tone: positive
Oct 7, 2026
Throughput
101.3 t/s gen · 1841.7 t/s pp
Quant
AWQ (AWQ)
VRAM reported
22 GB

Use cases

agenticlong-context

Summary

User reports Qwen3.6-27B-AWQ at 101.3 tok/s decode and 1841.7 tok/s prefill on dual modified RTX 2080 Ti 22GB cards with NVLink. Setup is vLLM 0.21.0 with AWQ Marlin, TP=2, MTP K=3, FlashInfer/FA2 attention, and FlashQLA SM70/SM75 legacy GDN prefill. The speed columns use the PP4096/TG128 repeat. The same rig reached a 735,084 token KV cache with turboquant_4bit_nc at max_model_len=262144, and passed a PP262000/TG1 gate at 785.26 tok/s prefill. A sequential 60-request Ragent6 run averaged 700.9 tok/s prefill and 35.2 tok/s generation. Gemma4 31B GPTQ reached 99.64 tok/s decode and 1655.65 tok/s prefill on the same runtime.