llamaperf

NVIDIA P102-100

NVIDIA · 10GB · 1 report

See what fits on this GPU →

Use the calculator to check which models fit in 10 GB of VRAM. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →
This page is thin (1 of 3 reports needed for indexing). Help fill it in.

Qwen3.6 35B (3B active)

NVIDIA P102-100 · llama.cpp · 32,768 ctx

Tone: positive
reported speed:
23.5 tokens/s generation · 432.3 tokens/s prompt processing
quant:
IQ4_XS (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports 70 t/s total across 3 concurrent users (23.3 t/s each) with 32K context per user. Uses two P102-100 cards (10GB each) for $100 total. Prompt processing speed 432 t/s. Model is Qwen3.6-35B-A3B at IQ4_XS quantization.