Qwen3.8 27B
on NVIDIA RTX 5060 Ti 16GB · llama.cpp · 262,144 ctx
Sep 29, 2026
Use cases
agentictool-uselong-contextvisioncoding
Summary
User reports Qwen3.8-27B at 256K context on an RTX 5060 Ti 16GB, with decode staying around 30-40 tok/s.
Setup is a KVMem-enabled llama.cpp server (v0.14.0) with ISTA IQ3_S quant, MTP, Q8 KV cache, vision on GPU, and a 32K GPU window plus 16K reserved for new tokens; the rest of the KV cache stays in 32GB system RAM.
In a 33-request tool task ending at 262058/262144 tokens, prefill was 437 tok/s first pass and 243 tok/s overall, decode 30 tok/s over the tool rounds and 29 tok/s on the last 512 tokens, with 15.5 GB peak VRAM and 13.1 GB RAM. An IQ4_XS variant with Q5 KV and CPU vision reached 466/255 tok/s prefill and 33/35 tok/s decode. A shorter image and code task ran about 37-41 tok/s decode.