llamaperf

Qwen3.8 27B

on NVIDIA RTX 2080 Ti 22GB (modded) · llama.cpp · 262,144 ctx

Tone: positive
Sep 28, 2026
Throughput
44.1 t/s gen
Quant
IQ3_S (GGUF)
KV cache
Q8_0
Flash Attention
on
System RAM
22 GB
VRAM reported
22 GB

Use cases

long-contextvision

Summary

User reports Qwen3.8-27B-GSQ-RCO at 44.07 t/s decode on a hardware-modded RTX 2080 Ti 22GB, with 38~44 t/s quoted as the overall range. Setup is a llama.cpp fork (TurboQuant 4-bit + Adaptive KV Streaming) with IQ3_S weights, 256K context (262,144 tokens), Q8_0 key cache and turbo4 value cache, a 2048 MiB GPU staging pool, and MTP speculative decoding with 2 draft tokens. Needle tests at 8K/16K/32K gave 41.23, 39.00 and 34.84 t/s decode with 81-83% draft acceptance and 100% recall at 82% depth; a sustained 256K run gave 40.55 t/s at 89.17% acceptance. Peak VRAM was 18,619 MiB and GPU power peaked at 266.4 W.