llamaperf
Sep 27, 2026
Throughput
139.1 t/s gen
Quant
NVFP4 (NVFP4)
System RAM
117 GB
VRAM reported
117 GB

Summary

User reports Qwen3.6-35B-A3B-NVFP4 at 139.1 tok/s on a single NVIDIA Jetson AGX Thor with 117 GB unified memory. Setup is vLLM built from source for sm_110a with DFlash speculative decoding (12 tokens), marlin MoE backend, flash_attn attention backend, 65536 context length, and 0.78 GPU memory utilization. The post also benchmarks Qwen3.5-4B-NVFP4 at 155.8 tok/s, Qwen3.6-27B-NVFP4 at 50.1 tok/s, and Qwen3.5-122B-A10B-NVFP4 at 52.6 tok/s, all at concurrency 1. The 122B requires cutlass MoE and TRITON_ATTN due to a Marlin crash at 256 experts, and achieves 27-42 tok/s with DFlash versus 10.9 tok/s autoregressive.