Qwen3.8 27B
on NVIDIA RTX 5060 Ti 16GB · llama.cpp · 110,000 ctx
Sep 17, 2026
Use cases
codingagentic
Summary
User reports Qwen3.8 27B at ~5.3 t/s on an RTX 5060 Ti 16GB with 16GB single-channel system RAM.
Setup is llama.cpp with IQ4 weights, 110k context, Q8 KV cache, and partial GPU/CPU offload; CPU and GPU each sit around 50% utilization.
User estimates Q8 weights plus 256k context would need 38-40GB total and drop to roughly 2 t/s, and asks whether quant level, context, or throughput should be prioritized for a local coding-agent worker.