Qwen3.8 27B
2× NVIDIA Tesla P100 16GB · llama.cpp
- reported speed:
- 32.6 tokens/s generation · 493.0 tokens/s prompt processing
- quant:
- Q6_K (GGUF)
- kv:
- q4_0
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3.8-27B Q6_K at 32.6 t/s decode (tg256, no MTP) on 2x Tesla P100 16GB with tensor split, versus 17.51 t/s upstream at the fork point. Setup is a llama.cpp fork with CUDA work for Pascal (sm_60), q4_0 KV cache, and fp16 math with fp32 accumulation. Prefill pp2048 at 0 context is 493 t/s versus ~250 t/s upstream. With MTP speculative decoding the fork reaches 54 t/s at 2k context and 29-35 t/s at 260k context. Prefill at 260k context is 123 t/s filling and 153 t/s for a question on a loaded context. Perplexity on the gate corpus at -c 4096 is 2.6101.