llamaperf

A100 80GB

NVIDIA · 80GB · 1 report

See what fits on this GPU →

Use the calculator to check which models fit in 80 GB of VRAM. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →
This page is thin (1 of 3 reports needed for indexing). Help fill it in.
reported speed:
56.8 tokens/s generation · 1019.1 tokens/s prompt processing
quant:
UD-Q2_K_XL (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks TensorSharp against llama.cpp on Qwen 3.8 Flash Next, reporting prompt processing and generation speeds for 1 GPU and 2 GPU layer split configurations. The promptTps and generationTps fields capture the 1 GPU pp512 and tg64 values respectively.

Sep 7, 2026