llamaperf

Qwen3.8 27B

on NVIDIA RTX 5060 Ti 16GB · llama.cpp · 4,096 ctx

Oct 3, 2026
Throughput
29.2 t/s gen
Quant
UD-IQ3_XXS (GGUF)
KV cache
q4_0
VRAM reported
16 GB

Use cases

coding

Summary

User reports Qwen 3.8 27B at 29.18 tok/s decode on an RTX 5060 Ti 16 GB at 4K context. Setup is llama.cpp build ad1de39e0 with UD-IQ3_XXS quant and q4_0 K/V cache, full GPU offload, flash attention, one parallel slot. Baseline decode without speculation is 29.18 tok/s at 4K, 23.62 at 32K and 19.77 at 64K; MTP n=2 gives 56.91/42.53/34.95 and MTP n=3 gives 63.89/45.48/39.06 tok/s at the same contexts. N-gram speculation added only 1.7% (29.66 tok/s). A 96K prompt measured 617.73 prompt tok/s and 36.74 decode tok/s.