llamaperf

RTX 4080

NVIDIA · 16GB · 3 reports

See what fits on this GPU →

Use the calculator to check which models fit in 16 GB of VRAM. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →

Qwen3.8 27B

RTX 4080 · ExLlamaV3 · 131,072 ctx

Tone: positive
reported speed:
56.5 tokens/s generation · 980.2 tokens/s prompt processing
quant:
SC_3.00bpw_H4_V4 (EXL3)
kv:
6,5
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 27B at 56.48 t/s decode and 980.2 t/s prefill on an RTX 4080 16GB at 131,072 context with 102,400 active input tokens. Setup is ExLlamaV3 1.5.0 with TabbyAPI, the turboderp SC_3.00bpw_H4_V4 EXL3 quant, a 6,5 KV cache, and MTP k=2 with a Q6 draft cache. VRAM peaked around 15.2GB of 16.4GB. Without MTP decode was 33.68 t/s, so MTP gave about a 68% decode increase at roughly a 5% prefill cost. Q4 draft cache was about 9% slower than Q6, and dynamic drafting was slower than fixed k=2.

Sep 14, 2026

Qwen3.8 27B

RTX 4080 · llama.cpp · 140,000 ctx

reported speed:
19.0 tokens/s generation
quant:
IQ4_XS (GGUF)
kv:
kvarn4
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports 8 t/s without MTP, 14 t/s with MTP, and about 19 t/s with full optimizations. Setup is beellama 0.4.3, a llama.cpp fork.

Sep 7, 2026
Tone: positive
reported speed:
8.0 tokens/s generation
quant:
Q4_K_M (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports running a 182B model on an RTX 4080 with 64GB RAM by offloading ngrams to SSD. User claims it is faster and more intelligent than Qwen3.8 27B and provides the launch command.

Sep 7, 2026