llamaperf

Qwen3.8 27B

on NVIDIA DGX Spark · SGLang · 262,144 ctx

Tone: positive
Oct 3, 2026
Throughput
41.0 t/s gen
Quant
NVFP4 (NVFP4)
KV cache
fp8_e4m3
System RAM
128 GB

Use cases

codingmathagentic

Summary

User benchmarks Qwen3.8-27B in NVFP4 on a single NVIDIA DGX Spark (GB10, 128 GB unified memory), reporting a greedy median of 41.0 tok/s with SGLang + DSpark speculative decoding. Setup is SGLang with NVFP4 weights, fp8_e4m3 KV cache, 262144 max context, and DSpark block-drafter speculation; vLLM + MTP is compared on the same harness. Per-workload SGLang figures range from 37.4-41.8 tok/s on code, 40.1-44.5 on reasoning, 40.3-52.9 on math, and 18.0-19.6 on free prose. vLLM single-stream reaches 20.0 tok/s on random 512/512 and 18.0 on ShareGPT; both engines converge to roughly 60 tok/s aggregate at concurrency 4. SGLang serves 4x the context (262K vs 65K) and has lower TTFT and TPOT.