llamaperf
Oct 4, 2026
Throughput
36.7 t/s gen · 899.0 t/s pp
Quant
NVFP4 (NVFP4)
KV cache
int8
System RAM
128 GB

Summary

User reports Qwen3.8-Flash-Next NVFP4 at 36.7 tok/s one-user decode and 899 tok/s prefill at 32k on a Jetson AGX Thor 128GB. Setup is TensorFold 0.6.0 with int8 KV cache, full 262,144-token window, --parallel 4. Prefill measured at 926 / 899 / 798 tok/s for 8k / 32k / 128k tokens; combined decode at 1 / 2 / 3 / 4 users is 34.8 / 53.7 / 70.3 / 76.7 tok/s; first answer to a 60k-token prompt takes 69.7 s. User compares against vLLM's FP8 build on the same box (38.4 tok/s decode, 2,101 tok/s prefill at 32k) and describes a local Thor prompt path with gated runs r09 and r11 reaching 1,549 tok/s prefill at 32k. 128 GB and 273 GB/s are published Thor specs, not re-measured.