llamaperf

Qwen3-VL 8B

on Jetson AGX Thor 64GB · TensorRT-Edge-LLM · 8,192 ctx

Oct 8, 2026
Throughput
88-92 t/s gen
Quant
NVFP4 (NVFP4)
System RAM
64 GB

Use cases

visionagenticlong-context

Summary

User reports Qwen3-VL-8B-Instruct at ~88-92 tok/s on an NVIDIA Jetson AGX Thor with the fast/tight 8K KV config, and ~70-78 tok/s with the 128K KV long-context config. Setup is a single Jetson AGX Thor with 64GB unified memory, TensorRT-Edge-LLM, NVFP4 weights and EAGLE-3 speculative decoding. EAGLE-3 averages ~3.76 accepted tokens per verify pass; cold start per request is ~6-8s and only one inference runs at a time.