Qwen3.8 27B
on NVIDIA DGX Spark · SGLang
Sep 30, 2026
Summary
User reports Qwen3.8-27B at about 34 tok/s real-world on a single DGX Spark (ASUS Ascent GX10, GB10, 128GB unified memory).
Setup is SGLang with an NVFP4 W4A4 checkpoint and DSpark block-speculative decoding, batch-1 single-stream.
User also measured 38.0 tok/s average on eval-style workloads and 46.7 tok/s peak on GSM8K-style prompts. On the same machine, llama.cpp UD-Q4_K_XL with MTP reached about 27 tok/s real-world and 24-30 tok/s on eval-style workloads, while vLLM 0.27 NVFP4 with MTP reached about 24.5 tok/s.