llamaperf
Sep 30, 2026
Throughput
34.0 t/s gen
Quant
NVFP4 (NVFP4)
System RAM
128 GB

Summary

User reports Qwen3.8-27B at about 34 tok/s real-world on a single DGX Spark (ASUS Ascent GX10, GB10, 128GB unified memory). Setup is SGLang with an NVFP4 W4A4 checkpoint and DSpark block-speculative decoding, batch-1 single-stream. User also measured 38.0 tok/s average on eval-style workloads and 46.7 tok/s peak on GSM8K-style prompts. On the same machine, llama.cpp UD-Q4_K_XL with MTP reached about 27 tok/s real-world and 24-30 tok/s on eval-style workloads, while vLLM 0.27 NVFP4 with MTP reached about 24.5 tok/s.