llamaperf

Tesla P40 24GB

NVIDIA · 24GB · 3 reports

See what fits on this GPU →

Use the calculator to check which models fit in 24 GB of VRAM. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →

Qwen3.8 27B

Tesla P40 24GB · 150,000 ctx

Tone: mixed
reported speed:
45.0 tokens/s generation · 450.0 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 27B at up to 45 t/s generation and 450 t/s prefill on 2x Tesla P40 with fresh context. At 150K+ context prefill falls to around 120 t/s and generation to 12-16 t/s. User notes that running two agents concurrently breaks prefix caching and KV cache sharing, and asks whether a single-supervisor harness with sequential subagent handoffs exists.

Sep 18, 2026
Tone: positive
reported speed:
91.0 tokens/s generation
quant:
Q8_K_XL (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen 3.8 27B at up to 91 t/s with MTP on 3x RTX 3090 plus a Tesla P40 and 128 GB of system RAM. Setup is the UD-Q8_K_XL quant. Thinking time varies by setting: low about 3s, medium about 3min, xHigh about 15min. The user compares it to Qwen 3.6 27B and DeepSeek V4 Flash, and says Qwen 3.8 produced a much better Galaga recreation with detailed sprites, sound effects, and a capture system. The user notes the high thinking time may be an issue for slower setups.

Sep 7, 2026
Tone: positive
reported speed:
48.0 tokens/s generation · 444.4 tokens/s prompt processing
quant:
Q8_0 (gguf)
kv:
F16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports up to 48 t/s with MTP speculative decoding on dual P40s, up from ~15 t/s. Long-context runs reach ~20 t/s, and prefill is 444 t/s at pp512. The model is a fine-tune of Qwen3.8 27B.

Sep 7, 2026