llamaperf
Oct 4, 2026
Throughput
25.0 t/s gen · 3557.0 t/s pp
Quant
NVFP4 (NVFP4)
KV cache
8-bit
System RAM
128 GB

Summary

User benchmarks Qwen3.8-27B at 24.96 t/s generation and 3557 t/s prompt processing on a Jetson AGX Thor 128GB, using the Mjolnir vLLM image with the FA4 GEMV decode kernel. Setup is vLLM 0.30.0 with NVFP4 weights and an 8-bit KV cache at 8K context, single request. The GEMV kernel is default-on and dispatches only on M=1, head_dim=256, GQA shapes. The GEMV leg wins at c=1 but regresses at c=4 (57.74 t/s at 8K context, −15.4% versus stock vLLM), which the user attributes to an open investigation. Stock vLLM reaches 24.43 t/s and 2468 t/s prompt at the same setting.