llamaperf
Oct 3, 2026
Throughput
24.5 t/s gen · 448.7 t/s pp
Quant
Q4_K_XL (GGUF)
KV cache
F16

Summary

User reports Qwen3.8-27B Q4_K_XL at 24.5 tok/s generation and 448.7 tok/s prompt processing on two Tesla P100 16GB cards with tensor split. Setup is a patched llama.cpp (upstream b10660 plus eleven patches) with F16 KV cache and full GPU offload across two cards. The figures are the after-patch numbers from a patch series that improves decode and prefill; the same run measured 22.3 tok/s generation and 427.0 tok/s prompt processing before the patches. A real 7,655-token request reached 398.5 tok/s prompt processing, and four concurrent agents reached 20.3 tok/s each (72.9 aggregate).