llamaperf

Qwen3.6 35B (3B active)

on 2× NVIDIA Tesla P100 16GB · llama.cpp

Tone: mixed
Sep 24, 2026
Throughput
85.0 t/s gen · 300.0 t/s pp
Quant
Q4_K_XL (GGUF)

Use cases

agentic

Summary

User reports Qwen3.6-35B-A3B at ~85 t/s generation and ~300 t/s prompt processing on 2x Tesla P100. Setup is llama.cpp with the Q4_K_XL GGUF, MTP speculative decoding, and community P100 patches that added about 50% decode; a single P100 on Q2_K_XL reached ~76 t/s generation and ~250 t/s prompt processing. Card count barely affects single-stream speed, and PCIe lane width (x16/x16, x16/x8, x8/x8) made no difference. The user corrects an earlier 3-card result that was invalidated by stuck 405 MHz core clocks.