llamaperf
Oct 6, 2026
Throughput
32.6 t/s gen · 493.0 t/s pp
Quant
Q6_K (GGUF)
KV cache
q4_0

Summary

User reports Qwen3.8-27B Q6_K at 32.6 t/s decode (tg256, no MTP) on 2x Tesla P100 16GB with tensor split, versus 17.51 t/s upstream at the fork point. Setup is a llama.cpp fork with CUDA work for Pascal (sm_60), q4_0 KV cache, and fp16 math with fp32 accumulation. Prefill pp2048 at 0 context is 493 t/s versus ~250 t/s upstream. With MTP speculative decoding the fork reaches 54 t/s at 2k context and 29-35 t/s at 260k context. Prefill at 260k context is 123 t/s filling and 153 t/s for a question on a loaded context. Perplexity on the gate corpus at -c 4096 is 2.6101.