Qwen3.8 27B
on NVIDIA RTX 5060 Ti 16GB · llama.cpp · 4,096 ctx
Oct 3, 2026
Use cases
coding
Summary
User reports Qwen 3.8 27B at 29.18 tok/s decode on an RTX 5060 Ti 16 GB at 4K context.
Setup is llama.cpp build ad1de39e0 with UD-IQ3_XXS quant and q4_0 K/V cache, full GPU offload, flash attention, one parallel slot.
Baseline decode without speculation is 29.18 tok/s at 4K, 23.62 at 32K and 19.77 at 64K; MTP n=2 gives 56.91/42.53/34.95 and MTP n=3 gives 63.89/45.48/39.06 tok/s at the same contexts. N-gram speculation added only 1.7% (29.66 tok/s). A 96K prompt measured 617.73 prompt tok/s and 36.74 decode tok/s.