Qwen3.6 35B (3B active)
on 2× NVIDIA Tesla P100 16GB · llama.cpp
Sep 24, 2026
Use cases
agentic
Summary
User reports Qwen3.6-35B-A3B at ~85 t/s generation and ~300 t/s prompt processing on 2x Tesla P100.
Setup is llama.cpp with the Q4_K_XL GGUF, MTP speculative decoding, and community P100 patches that added about 50% decode; a single P100 on Q2_K_XL reached ~76 t/s generation and ~250 t/s prompt processing.
Card count barely affects single-stream speed, and PCIe lane width (x16/x16, x16/x8, x8/x8) made no difference. The user corrects an earlier 3-card result that was invalidated by stuck 405 MHz core clocks.