llamaperf
Oct 7, 2026
Throughput
95.5 t/s gen · 465.8 t/s pp
Quant
Q4_K_P (GGUF)
KV cache
Q8_0
System RAM
30 GB
VRAM reported
30 GB

Use cases

agenticlong-context

Summary

User reports Qwen3.6-35B-A3B at 95.5 t/s decode and 465.8 t/s prefill on a mixed three-GPU Turing setup of 2x CMP 50HX 10GB plus 1x RTX 2080 Ti 22GB. Setup is llama.cpp with a ported DP2A patch (PR #25834) and -fmad=false, Q4_K_P GGUF weights, Q8_0 KV cache, Flash Attention, 262144 context, MTP speculative decoding at --spec-draft-n-max 3, tensor split 1,1,2.5 with the RTX as tail stage, ubatch 448. Baseline DP4A gave 47.8 t/s decode and 372.7 t/s prefill; DP2A alone 54.0 t/s; adding -fmad=false 62.2 t/s. A repo-28k workload measured 405.4 t/s prefill and 82.3 t/s decode. An end-to-end rerun showed 91.4 t/s raw eval, 4% below the 95.5 reference. Context 368640 also works but quality beyond the trained window was not measured.