Qwen3.8 27B
on NVIDIA RTX 2080 Ti 22GB (modded) · KVMem · 262,144 ctx
Oct 7, 2026
Use cases
long-contextagentictool-usevision
Summary
User reports Qwen3.8-27B at 36.6 tok/s decode on a single modded RTX 2080 Ti 22GB, with 262,144 tokens of context.
Setup is KVMem (retrieval-based long context) with IQ3_S weights, q8_0 KV cache, MTP speculative decoding (58.6%/61.6% acceptance) and vision, using 16,552 of 22,528 MiB VRAM.
The user compares against the upstream author's RTX 5060 Ti 16GB at 31.7 tok/s, and notes the lossless KV-streaming alternative runs 30-42 tok/s in the resident window but drops to ~9.6 tok/s past 135K tokens; a 260,096-token needle-in-a-haystack test hit exactly.