llamaperf

V100 16GB

NVIDIA · 16GB · 2 reports

See what fits on this GPU →

Use the calculator to check which models fit in 16 GB of VRAM. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →
This page is thin (2 of 3 reports needed for indexing). Help fill it in.

Qwen3.8 27B

V100 16GB · 256,000 ctx

Tone: positive
reported speed:
35.0 tokens/s generation · 650.0 tokens/s prompt processing
quant:
Q8
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticlong-context

User reports Qwen3.8 27B at 30-40 t/s decode and 600-700 t/s prefill on 3x V100 16GB. Setup uses Q8 quantization with speculative decoding, MTP, and prefix caching, tuned by an automated agent. On heavy agentic work at 256k context, decode dropped to around 20 t/s. Qwen3.8 Flash Next at Q4 ran at 20 t/s decode and 90 t/s prefill. The build cost around $1500 and uses a 3D printed cooling block; GPU temperatures stay under 55C.

Sep 13, 2026

Qwen3.8 27B

V100 16GB · v100-skinny

Tone: positive
reported speed:
219.1 tokens/s generation
quant:
NVFP4
kv:
F16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 NVFP4 on four Tesla V100s (2017) via a custom v100-skinny engine matching RTX 5090 decode throughput at ~220 tok/s. Setup uses QPN kernels to translate FP4/FP8 to FP16 for Volta tensor cores. Long-context tests show MTP depth k=3 better at ~65K context. The figure is for 4 GPUs versus 1, is not power efficient, and prefill is slower.

Aug 28, 2026