llamaperf

NVIDIA RTX 3080 10GB

NVIDIA · 10GB · 2 reports

Engines people use on it: llama.cpp 1

Run models on your NVIDIA RTX 3080 10GB? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the NVIDIA RTX 3080 10GB

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 10 GB of VRAM.

This page is thin (2 of 3 reports needed for indexing). Help fill it in.

Qwen3.8 27B

NVIDIA RTX 3080 10GB · Unsloth · 44,000 ctx

reported speed:
20-25 tokens/s generation
quant:
IQ3_XXS (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticcoding

User asks whether replacing a GTX 1070 with an RTX 3060 12GB is worthwhile for local LLMs and agentic coding. Current setup is an RTX 3080 10GB plus GTX 1070 8GB, Ryzen 5 5600X, 32GB RAM. User reports running Qwen3.8 27B IQ3_XXS at about 20-25 tokens/s with 44K context using Unsloth, and wants at least 128K context. The 3060 upgrade is a purchase question, not a measured run; no throughput figure is reported for the proposed card.

Oct 5, 2026

Bonsai 2 27B

NVIDIA RTX 3080 10GB · llama.cpp · 32,768 ctx

reported speed:
52.2 tokens/s generation · 1210.6 tokens/s prompt processing
quant:
PTQ1_0 (GGUF)
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Bonsai 2 27B PTQ1_0 at 52.18 t/s generation and 1210.6 t/s prefill on an RTX 3080 10GB. Setup is llama.cpp (fork build prism-b10735-842b188, CUDA 12.8) with PTQ1_0 quant and q8_0 KV cache at 32K context, single card, batch 1, depth 0, -ngl 99 -fa on. User also benchmarks PQ2_0 at 61.38 t/s and Qwen3.8-27B-UD-IQ2_XXS at 44.46 t/s on the same card, and reports a context ladder up to 160K q4_0 at 9627 MiB peak VRAM. A LiveCodeBench v6 comparison gives PTQ1_0 32/50 pass@1 versus 16/50 for IQ2_XXS and 28/50 for PQ2_0.

Sep 28, 2026

Get a weekly email of new NVIDIA RTX 3080 10GB reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Bonsai 2 27B
NVIDIA RTX 3080 10GB
PTQ1_0
llama.cpp
32,76852.2 tokens/s