Qwen3.8 27B
NVIDIA RTX 5060 8GB · llama.cpp · 8,192 ctx
- reported speed:
- 30.1 tokens/s generation
- quant:
- UD-IQ2_XXS (GGUF)
- kv:
- q4_0
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3.8-27B at 30.10 tok/s median decode on an RTX 5060 8GB, over five runs at an occupied 8192-token prompt. Setup is llama.cpp with UD-IQ2_XXS weights and q4_0 KV cache, 65/65 layers on CUDA, peak VRAM 7767 MiB. A short-context FULL_GPU decode of 31.39 tok/s and a ctx512 control of 18.11 tok/s are also given; quality gate and uncensored checkpoint are not done.