llamaperf

Qwen3.8 27B

on NVIDIA RTX 5060 Ti 16GB · llama.cpp · 110,000 ctx

Sep 17, 2026
Throughput
5.3 t/s gen
Quant
IQ4 (GGUF)
KV cache
Q8
System RAM
16 GB
VRAM reported
16 GB

Use cases

codingagentic

Summary

User reports Qwen3.8 27B at ~5.3 t/s on an RTX 5060 Ti 16GB with 16GB single-channel system RAM. Setup is llama.cpp with IQ4 weights, 110k context, Q8 KV cache, and partial GPU/CPU offload; CPU and GPU each sit around 50% utilization. User estimates Q8 weights plus 256k context would need 38-40GB total and drop to roughly 2 t/s, and asks whether quant level, context, or throughput should be prioritized for a local coding-agent worker.