Qwen3.8 27B
on NVIDIA RTX 5060 Ti 16GB · llama.cpp · 94,208 ctx
Oct 3, 2026
Summary
User reports Qwen3.8-27B at a steady 50-55 t/s with MTP on a single RTX 5060 Ti 16GB.
Setup is llama.cpp with UD-IQ3_XXS GGUF weights and Q4_0 KV cache for K and V, 94208-token context, batch and ubatch 512, flash attention on, draft-mtp speculation with spec-draft-n-max 3, parallel 1.
Without MTP the average is 35 t/s. The machine has an Intel i5 10th gen and 32GB DDR4.