Qwen3.8 27B
on 2× NVIDIA Tesla P100 16GB · llama.cpp
Oct 6, 2026
Summary
User reports Qwen3.8-27B Q6_K at 32.6 t/s decode (tg256, no MTP) on 2x Tesla P100 16GB with tensor split, versus 17.51 t/s upstream at the fork point.
Setup is a llama.cpp fork with CUDA work for Pascal (sm_60), q4_0 KV cache, and fp16 math with fp32 accumulation. Prefill pp2048 at 0 context is 493 t/s versus ~250 t/s upstream.
With MTP speculative decoding the fork reaches 54 t/s at 2k context and 29-35 t/s at 260k context. Prefill at 260k context is 123 t/s filling and 153 t/s for a question on a loaded context. Perplexity on the gate corpus at -c 4096 is 2.6101.