llamaperf

Qwen3.8 27B

on NVIDIA RTX 5060 Ti 16GB · llama.cpp · 262,144 ctx

Tone: positive
Sep 29, 2026
Throughput
30-40 t/s gen
Quant
IQ3_S (GGUF)
KV cache
Q8
System RAM
32 GB
VRAM reported
16 GB

Use cases

agentictool-uselong-contextvisioncoding

Summary

User reports Qwen3.8-27B at 256K context on an RTX 5060 Ti 16GB, with decode staying around 30-40 tok/s. Setup is a KVMem-enabled llama.cpp server (v0.14.0) with ISTA IQ3_S quant, MTP, Q8 KV cache, vision on GPU, and a 32K GPU window plus 16K reserved for new tokens; the rest of the KV cache stays in 32GB system RAM. In a 33-request tool task ending at 262058/262144 tokens, prefill was 437 tok/s first pass and 243 tok/s overall, decode 30 tok/s over the tool rounds and 29 tok/s on the last 512 tokens, with 15.5 GB peak VRAM and 13.1 GB RAM. An IQ4_XS variant with Q5 KV and CPU vision reached 466/255 tok/s prefill and 33/35 tok/s decode. A shorter image and code task ran about 37-41 tok/s decode.