llamaperf

Qwen3.8 27B

on NVIDIA RTX 5060 8GB · llama.cpp · 8,192 ctx

Oct 7, 2026
Throughput
30.1 t/s gen
Quant
UD-IQ2_XXS (GGUF)
KV cache
q4_0
VRAM reported
8 GB

Summary

User reports Qwen3.8-27B at 30.10 tok/s median decode on an RTX 5060 8GB, over five runs at an occupied 8192-token prompt. Setup is llama.cpp with UD-IQ2_XXS weights and q4_0 KV cache, 65/65 layers on CUDA, peak VRAM 7767 MiB. A short-context FULL_GPU decode of 31.39 tok/s and a ctx512 control of 18.11 tok/s are also given; quality gate and uncensored checkpoint are not done.