Qwen3.6 35B (3B active)
on NVIDIA Tesla P100 16GB · llama.cpp · 32,768 ctx
Sep 27, 2026
Use cases
codingcreative-writing
Summary
User reports Qwen3.6 35B A3B at 54-60 t/s generation (66-72 t/s on code) on a single Tesla P100 16GB, up from 30-35 t/s prose on an RX 6600 XT.
Setup is llama.cpp with shinbunbun patches, UD_Q4_K_XL quant, 16-bit KV cache, MTP speculative decoding, and --n-cpu-moe 22, running 32k context.
Prefill is 600 t/s at 0 ctx dropping to 440-500 t/s by 10k. The user notes the P100 is underrated for the price and has a second card coming for full offload.