llamaperf

Qwen3.8 27B

on NVIDIA RTX 5060 Ti 16GB · llama.cpp · 100,000 ctx

Tone: positive
Sep 28, 2026
Throughput
25-30 t/s gen
Quant
IQ3_XXS (GGUF)
KV cache
q8_0
VRAM reported
16 GB

Use cases

visionlong-context

Summary

User reports Qwen3.8 27B running on an RTX 5060 Ti 16GB with MTP speculative decoding, reaching up to 50 t/s in the first half of a 100k context and dropping to about 25-30 t/s at the end. Setup is a modified llama.cpp fork with adaptive KV cache streaming, IQ3_XXS GGUF weights, q8_0 K cache and q4_0 V cache, a 2100 MiB KV stream buffer, and a BF16 vision head. The user notes the streaming fork only works on NVIDIA hardware and that the speed depends on context size versus KV stream buffer size.