Qwen3.8 27B
on NVIDIA RTX 5060 8GB · llama.cpp · 8,192 ctx
Oct 7, 2026
Summary
User reports Qwen3.8-27B at 30.10 tok/s median decode on an RTX 5060 8GB, over five runs at an occupied 8192-token prompt.
Setup is llama.cpp with UD-IQ2_XXS weights and q4_0 KV cache, 65/65 layers on CUDA, peak VRAM 7767 MiB.
A short-context FULL_GPU decode of 31.39 tok/s and a ctx512 control of 18.11 tok/s are also given; quality gate and uncensored checkpoint are not done.