llamaperf

Qwen3.8 125B (6B active) Flash-Next

on NVIDIA DGX Spark · SGLang · 262,144 ctx

Tone: positive
Sep 18, 2026
Throughput
35.0 t/s gen
Quant
NVFP4
System RAM
128 GB

Use cases

codingagentic

Summary

User reports Qwen3.8 Flash-Next at about 35 t/s on a single DGX Spark. Setup is SGLang with an NVFP4 quant at 262,144 context, used through VSCode Copilot in autopilot mode. The session produced about 10k lines of code and consumed about 800k tokens over 8 hours of planning, coding, and testing.