llamaperf
Sep 27, 2026
Throughput
101.3 t/s gen · 3191.0 t/s pp
Quant
W4A16 (W4A16)
KV cache
BF16
System RAM
92 GB
VRAM reported
64 GB

Use cases

codinglong-contexttool-usemultilingual

Summary

User reports Qwen3.8-Flash-Next at 101.3 t/s decode on Japanese prose and 170.0 t/s on code, single stream, on two NVIDIA CMP 170HX cards. Setup is vLLM with W4A16 weights, BF16 KV cache, MTP k=4 speculative decoding, 262,144-token context, and expert parallel across the two cards. The PLE n-gram table was converted to FP8 locally to fit in 92 GiB of host RAM. Aggregate throughput reaches 392 t/s at 4 concurrent requests. Prefill measures 3,191 t/s at 6,954 tokens and 3,261 t/s at 27,853 tokens. A 200,087-token prompt returned the planted code with 70.5 s time to first token and 69.0 t/s decode at depth. MTP is worth about 1.6x on prose and 2.6x on code. The user notes xhigh reasoning effort spent the whole budget without answering in 5 of 6 runs.