llamaperf

NVIDIA Tesla P100 16GB

NVIDIA · 16GB · 8 reports

Engines people use on it: llama.cpp 6

Run models on your NVIDIA Tesla P100 16GB? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the NVIDIA Tesla P100 16GB

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 16 GB of VRAM.

reported speed:
32.6 tokens/s generation · 493.0 tokens/s prompt processing
quant:
Q6_K (GGUF)
kv:
q4_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8-27B Q6_K at 32.6 t/s decode (tg256, no MTP) on 2x Tesla P100 16GB with tensor split, versus 17.51 t/s upstream at the fork point. Setup is a llama.cpp fork with CUDA work for Pascal (sm_60), q4_0 KV cache, and fp16 math with fp32 accumulation. Prefill pp2048 at 0 context is 493 t/s versus ~250 t/s upstream. With MTP speculative decoding the fork reaches 54 t/s at 2k context and 29-35 t/s at 260k context. Prefill at 260k context is 123 t/s filling and 153 t/s for a question on a loaded context. Perplexity on the gate corpus at -c 4096 is 2.6101.

Oct 6, 2026
Tone: positive
reported speed:
110.0 tokens/s generation
quant:
Q4_0 (GGUF)
kv:
F16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagenticlong-context

User reports Nemotron-3.5-Lightning-30B-A3B at 110 tok/s writing code on two Tesla P100 16GB cards. Setup is llama.cpp b10970 with Q4_0 weights, F16 KV cache, tensor split across both cards, and the model's built-in MTP draft head at n-max 2. The same model reaches 95 tok/s on prose, 50 tok/s at 128k context, 39 tok/s at 256k, and 16 tok/s at 1M tokens. The study covers 545 speed measurements of 29 models from 2B to 122B parameters, all weights and KV cache in VRAM with no system RAM offload. A 119B MoE model runs 39 tok/s against 4.3 tok/s for a 70B dense model on the same cards. Tensor split makes dense models from 8B up 21-44% faster. A single P100 throttles to 906 MHz and loses 25% under sustained load, while two cards share the heat and lose 5.5%.

Oct 6, 2026
reported speed:
24.5 tokens/s generation · 448.7 tokens/s prompt processing
quant:
Q4_K_XL (GGUF)
kv:
F16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8-27B Q4_K_XL at 24.5 tok/s generation and 448.7 tok/s prompt processing on two Tesla P100 16GB cards with tensor split. Setup is a patched llama.cpp (upstream b10660 plus eleven patches) with F16 KV cache and full GPU offload across two cards. The figures are the after-patch numbers from a patch series that improves decode and prefill; the same run measured 22.3 tok/s generation and 427.0 tok/s prompt processing before the patches. A real 7,655-token request reached 398.5 tok/s prompt processing, and four concurrent agents reached 20.3 tok/s each (72.9 aggregate).

Oct 3, 2026

Qwen3.8 27B

2× NVIDIA Tesla P100 16GB · llama.cpp · 262,144 ctx

Tone: positive
reported speed:
50-60 tokens/s generation · 350.0 tokens/s prompt processing
quant:
Q6_K (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingcreative-writing

User reports Qwen3.8 27B at 50-60 t/s generation and 350 t/s prefill on 2x Tesla P100 16GB with a custom llama.cpp fork. Setup is llama.cpp with Q6_K quant, 262144 context, batch 32768, ubatch 1024, and MTP speculative decoding. At 260k context, generation is 30-35 t/s and prefill is 110 t/s. GPUs are capped at 175W/250W each and run at 79C with minor thermal throttling; user estimates 5-10% higher numbers with better cooling.

Sep 27, 2026
Tone: positive
reported speed:
54-60 tokens/s generation · 440-500 tokens/s prompt processing
quant:
UD_Q4_K_XL (GGUF)
kv:
16bit

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingcreative-writing

User reports Qwen3.6 35B A3B at 54-60 t/s generation (66-72 t/s on code) on a single Tesla P100 16GB, up from 30-35 t/s prose on an RX 6600 XT. Setup is llama.cpp with shinbunbun patches, UD_Q4_K_XL quant, 16-bit KV cache, MTP speculative decoding, and --n-cpu-moe 22, running 32k context. Prefill is 600 t/s at 0 ctx dropping to 440-500 t/s by 10k. The user notes the P100 is underrated for the price and has a second card coming for full offload.

Sep 27, 2026
Tone: mixed
reported speed:
85.0 tokens/s generation · 300.0 tokens/s prompt processing
quant:
Q4_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agentic

User reports Qwen3.6-35B-A3B at ~85 t/s generation and ~300 t/s prompt processing on 2x Tesla P100. Setup is llama.cpp with the Q4_K_XL GGUF, MTP speculative decoding, and community P100 patches that added about 50% decode; a single P100 on Q2_K_XL reached ~76 t/s generation and ~250 t/s prompt processing. Card count barely affects single-stream speed, and PCIe lane width (x16/x16, x16/x8, x8/x8) made no difference. The user corrects an earlier 3-card result that was invalidated by stuck 405 MHz core clocks.

Sep 24, 2026
Tone: positive
reported speed:
70.0 tokens/s generation
quant:
Q4_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports almost 70 t/s with Qwen3.6-35B-A3B Q4_K_XL on a budget build using three Nvidia Tesla P100 16GB GPUs. The GPUs cost $80 each and are split across three nodes to manage thermals; the motherboard required a patched BIOS to enable Above 4G Decoding. User is still testing and optimizing, and has published a GitHub repository for the build.

Sep 22, 2026
Tone: positive
reported speed:
40.0 tokens/s generation
quant:
Q6_K (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 27B at 40+ t/s on two Tesla P100s, with a roughly 30 t/s average across long contexts and up to 55 t/s at 0 context. Setup is a custom llama.cpp fork with P100 kernel optimizations, Q6_K quant, both cards capped at 175W and communicating over PCIe gen 3. Speeds vary by about ±2 t/s; the user notes fp16 math saves 40% or more on prefill with negligible accuracy loss.

Sep 20, 2026

Get a weekly email of new NVIDIA Tesla P100 16GB reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Qwen3.8 27B
2× NVIDIA Tesla P100 16GB
Q6_K
llama.cpp
Not reported32.6 tokens/s
Nemotron 3.5 Lightning 30B (3B active)
2× NVIDIA Tesla P100 16GB
Q4_0
llama.cpp
Not reported110.0 tokens/s
Qwen3.8 27B
2× NVIDIA Tesla P100 16GB
Q4_K_XL
llama.cpp
Not reported24.5 tokens/s
Qwen3.6 35B (3B active)
2× NVIDIA Tesla P100 16GB
Q4_K_XL
llama.cpp
Not reported85.0 tokens/s
Qwen3.6 35B (3B active)
3× NVIDIA Tesla P100 16GB
Q4_K_XL
Engine not reported
Not reported70.0 tokens/s
Qwen3.8 27B
2× NVIDIA Tesla P100 16GB
Q6_K
llama.cpp
Not reported40.0 tokens/s