llamaperf
Oct 5, 2026
Throughput
1680.0 t/s pp
Quant
NVFP4 (GGUF)
KV cache
INT8
System RAM
64 GB
VRAM reported
32 GB

Summary

User reports Qwen3.8-Flash-Next NVFP4 running on a single RTX PRO 4500 Blackwell 32GB with 64GB DDR5 system RAM, reaching up to ~80 tok/s at ~50k context and ~60-67 tok/s at ~188k warm context. Setup is the Strata NVFP4 fork with NVFP4 routed experts (~63.3 GiB), separate FP8 PLE, INT8 KV cache, MTP speculative decoding, and W4A8 prefill on Blackwell, served through Strata's OpenAI-compatible server. A cold 189k full prompt measured ~1,680 tok/s prefill and ~53 tok/s decode. User calls the result impressive for a single 32GB GPU with only 64GB system RAM.