llamaperf

Qwen3.8 27B

on NVIDIA RTX 3080 10GB · Unsloth · 44,000 ctx

Oct 5, 2026
Throughput
20-25 t/s gen
Quant
IQ3_XXS (GGUF)
System RAM
32 GB

Use cases

agenticcoding

Summary

User asks whether replacing a GTX 1070 with an RTX 3060 12GB is worthwhile for local LLMs and agentic coding. Current setup is an RTX 3080 10GB plus GTX 1070 8GB, Ryzen 5 5600X, 32GB RAM. User reports running Qwen3.8 27B IQ3_XXS at about 20-25 tokens/s with 44K context using Unsloth, and wants at least 128K context. The 3060 upgrade is a purchase question, not a measured run; no throughput figure is reported for the proposed card.