llamaperf
Sep 18, 2026
Throughput
11352.0 t/s pp
Quant
NVFP4 (NVFP4)
System RAM
96 GB
VRAM reported
96 GB

Use cases

tool-uselong-contextagenticcreative-writing

Summary

User reports Qwen3.8-Flash-Next on one RTX PRO 6000 Blackwell 96GB, with the official SGLang NVFP4 image cutting time to first token at roughly 254K context from 34.9s to 22.4s and raising prefill from 7,284 to 11,352 tok/s. Setup is the lmsysorg/sglang:dev-qwen38-next-local image with the RadixArk NVFP4 checkpoint, a 262,144-token window and one slot; the previous build streamed the 47.68 GiB per-layer embedding table from NVMe while the official image pins it in host RAM. A separate cache test dropped first-token time from 21.7s cold to 0.44s cached. Long-context recall scored 81/81, BFCL single-turn tool accuracy 85.2% versus 49.5% multi-turn, and tau2-bench telecom completion 68.1% with 41.2% passing all three attempts.