llamaperf
Sep 27, 2026
Throughput
96.9 t/s gen
Quant
UD-Q4_K_XL (GGUF)
KV cache
q8_0
System RAM
15 GB
VRAM reported
160 GB

Use cases

long-contextagentic

Summary

User reports Qwen3.8-Flash-Next at 96.9 tok/s single-request decode on four NVIDIA CMP 170HX 40GB cards. Setup is llama.cpp with UD-Q4_K_XL weights, q8_0 KV cache, MTP speculative decoding (draft length 4, GPU sampling), 262K context, and the SM clock pinned at 1410 MHz. The same configuration reaches 87.7 tok/s at 70K context; three concurrent 3K requests give 34-37 tok/s each (about 100 tok/s aggregate). Earlier revisions of the fork measured 64-73 tok/s at 2K and 58-70 tok/s at 70K, against a baseline fork at 46-55 tok/s and 27-45 tok/s respectively.