llamaperf

RTX 4080 Super

NVIDIA · 16GB · 1 report

See what fits on this GPU →

Use the calculator to check which models fit in 16 GB of VRAM. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →
This page is thin (1 of 3 reports needed for indexing). Help fill it in.

Qwen3.8 27B Escha-W2

RTX 4080 Super · SGLang · 98,304 ctx

Tone: positive
reported speed:
59.0 tokens/s generation
quant:
2.469 bpw
kv:
FP8 E4M3

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagenticlong-context

User pushed Qwen3.8-27B-Escha-W2 to 98K context on a 16GB 4080 Super. Reports ~59 tok/s short context, ~50.5 tok/s at 60K context. Quality surprisingly good despite aggressive 2.469 bpw quant. MTP4 gave ~67.7 tok/s at 64K context but chose no speculation for max context. Setup uses Escha's SGLang build with FP8 KV cache and BF16 SSM state.