Qwen3.8 27B
on NVIDIA RTX 5060 Ti 16GB · llama.cpp · 100,000 ctx
Sep 28, 2026
Use cases
visionlong-context
Summary
User reports Qwen3.8 27B running on an RTX 5060 Ti 16GB with MTP speculative decoding, reaching up to 50 t/s in the first half of a 100k context and dropping to about 25-30 t/s at the end.
Setup is a modified llama.cpp fork with adaptive KV cache streaming, IQ3_XXS GGUF weights, q8_0 K cache and q4_0 V cache, a 2100 MiB KV stream buffer, and a BF16 vision head.
The user notes the streaming fork only works on NVIDIA hardware and that the speed depends on context size versus KV stream buffer size.