Qwen3.8 27B
NVIDIA RTX 3080 10GB · Unsloth · 44,000 ctx
- reported speed:
- 20-25 tokens/s generation
- quant:
- IQ3_XXS (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User asks whether replacing a GTX 1070 with an RTX 3060 12GB is worthwhile for local LLMs and agentic coding. Current setup is an RTX 3080 10GB plus GTX 1070 8GB, Ryzen 5 5600X, 32GB RAM. User reports running Qwen3.8 27B IQ3_XXS at about 20-25 tokens/s with 44K context using Unsloth, and wants at least 128K context. The 3060 upgrade is a purchase question, not a measured run; no throughput figure is reported for the proposed card.