llamaperf

Qwen3.8 27B

on NVIDIA RTX 2080 Ti 22GB (modded) · KVMem · 262,144 ctx

Tone: positive
Oct 7, 2026
Throughput
36.6 t/s gen
Quant
IQ3_S (GGUF)
KV cache
q8_0
VRAM reported
22 GB

Use cases

long-contextagentictool-usevision

Summary

User reports Qwen3.8-27B at 36.6 tok/s decode on a single modded RTX 2080 Ti 22GB, with 262,144 tokens of context. Setup is KVMem (retrieval-based long context) with IQ3_S weights, q8_0 KV cache, MTP speculative decoding (58.6%/61.6% acceptance) and vision, using 16,552 of 22,528 MiB VRAM. The user compares against the upstream author's RTX 5060 Ti 16GB at 31.7 tok/s, and notes the lossless KV-streaming alternative runs 30-42 tok/s in the resident window but drops to ~9.6 tok/s past 135K tokens; a 260,096-token needle-in-a-haystack test hit exactly.