llamaperf

Qwen3.8 27B Uncensored

on NVIDIA RTX 5060 Ti 16GB · llama.cpp · 131,072 ctx

Tone: mixed
Oct 1, 2026
Throughput
35.0 t/s gen
Quant
IQ3_XXS (GGUF)
KV cache
q4_0
System RAM
16 GB
VRAM reported
16 GB

Summary

User reports about 35 t/s or more with Qwen3.8 27B GSQ-RCO-Uncensored on an RTX 5060 Ti 16GB. Setup is llama.cpp with IQ3_XXS quant, q4_0 KV cache, 131072 context, and MTP draft plus ngram speculative decoding. The user is on Fedora 44 with an AMD Ryzen 9600x and 16 GB system RAM, and had struggled to get reasonable speed before finding this model.