llamaperf

Qwen3.8 27B

on NVIDIA RTX 5060 Ti 16GB · llama.cpp · 94,208 ctx

Oct 3, 2026
Throughput
50-55 t/s gen
Quant
UD-IQ3_XXS (GGUF)
KV cache
Q4_0
Flash Attention
on
System RAM
32 GB
VRAM reported
16 GB

Summary

User reports Qwen3.8-27B at a steady 50-55 t/s with MTP on a single RTX 5060 Ti 16GB. Setup is llama.cpp with UD-IQ3_XXS GGUF weights and Q4_0 KV cache for K and V, 94208-token context, batch and ubatch 512, flash attention on, draft-mtp speculation with spec-draft-n-max 3, parallel 1. Without MTP the average is 35 t/s. The machine has an Intel i5 10th gen and 32GB DDR4.