llamaperf

RTX A6000 48GB

NVIDIA · 48GB · 2 reports

See what fits on this GPU →

Use the calculator to check which models fit in 48 GB of VRAM. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →
This page is thin (2 of 3 reports needed for indexing). Help fill it in.
reported speed:
17.2 tokens/s generation · 70.0 tokens/s prompt processing
quant:
Q8_K_XL

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports prompt processing in the high 70s t/s, dropping to the mid 30s at 300k context. The full 1M context fits in 48 GB of VRAM, but prompt processing is slow.

Sep 7, 2026
reported speed:
16.9 tokens/s generation
quant:
bf16 (safetensors)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

text-generation

User reports bf16 with no quantization at 10.25 GB VRAM and 61 ms TTFT. Source is dev.to Gaurav Vij.

May 1, 2026