llamaperf

Qwen3-VL

1 report

Qwen3-VL VRAM requirements by size and quant →

How does Qwen3-VL run on your hardware?

Pick your GPU or Mac, then read the speeds people reported for Qwen3-VL on it, or estimate memory fit and speed.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Run Qwen3-VL yourself? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance for Qwen3-VL

Filter this model’s reports by setup →
Thin page (1 of 3 reports needed for indexing). Add yours.

Qwen3-VL 8B

Jetson AGX Thor 64GB · TensorRT-Edge-LLM · 8,192 ctx

reported speed:
88-92 tokens/s generation
quant:
NVFP4 (NVFP4)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

visionagenticlong-context

User reports Qwen3-VL-8B-Instruct at ~88-92 tok/s on an NVIDIA Jetson AGX Thor with the fast/tight 8K KV config, and ~70-78 tok/s with the 128K KV long-context config. Setup is a single Jetson AGX Thor with 64GB unified memory, TensorRT-Edge-LLM, NVFP4 weights and EAGLE-3 speculative decoding. EAGLE-3 averages ~3.76 accepted tokens per verify pass; cold start per request is ~6-8s and only one inference runs at a time.

Oct 8, 2026

Get a weekly email of new Qwen3-VL reports on any GPU.

Email me new reports