llamaperf

Local model performance on your hardware

Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.

What runs on your hardware?

Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Local model performance reports from the community

These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →

Model: Qwen3-VL
Compare setup details

Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.

Qwen3-VL 8B

Jetson AGX Thor 64GB · TensorRT-Edge-LLM · 8,192 ctx

reported speed:
88-92 tokens/s generation
quant:
NVFP4 (NVFP4)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

visionagenticlong-context

User reports Qwen3-VL-8B-Instruct at ~88-92 tok/s on an NVIDIA Jetson AGX Thor with the fast/tight 8K KV config, and ~70-78 tok/s with the 128K KV long-context config. Setup is a single Jetson AGX Thor with 64GB unified memory, TensorRT-Edge-LLM, NVFP4 weights and EAGLE-3 speculative decoding. EAGLE-3 averages ~3.76 accepted tokens per verify pass; cold start per request is ~6-8s and only one inference runs at a time.

Oct 8, 2026
Showing 1–1 of 1
Page 1 of 1

Community benchmarks snapshot

Records by GPU

NVIDIA RTX 3090172NVIDIA RTX 5090120AMD Strix Halo 128GB87NVIDIA DGX Spark68NVIDIA RTX 5060 Ti 16GB56NVIDIA RTX Pro 6000 Blackwell53NVIDIA RTX 3060 12GB48AMD Radeon AI PRO R9700 32GB45NVIDIA RTX 409038NVIDIA RTX 5070 Ti32

Records by model

1488 total
Qwen3.8820
Qwen3.6177
DeepSeek V4 Flash128
Gemma 466
Qwen3.537
Qwen326
other234

Records by engine

1141 total
llama.cpp593
vLLM166
Strata52
NInfer41
Ollama36
other253

Use cases

coding 483agentic 311long-context 228tool-use 127vision 92summarization 47math 37creative-writing 31multilingual 19text-generation 9rp 6reasoning 3
coding483agentic311long-context228tool-use127vision92summarization47math37creative-writing31

Median t/s by GPU

On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking

RTX 509082RTX Pro 600067M5 Ultra 256GB50RX 7900 XTX41RTX 309036RTX 5090 Laptop 24GB34V100 32GB33RTX 409032RX 7800 XT 16GB30Radeon AI PRO R9700 32GB29

Reports by model size

Qwen3.8 27B489Qwen3.8 125B · 6B active301Qwen3.6 35B · 3B active105DeepSeek V4 Flash 284B · 13B active102Qwen3.6 27B70Gemma 4 26B · 4B active27DeepSeek V4.1 Flash 552B · 16B active24GLM-5.3 320B · 18B active21

Quants

Q4_K_M127NVFP4104IQ4_XS61Q4_K_XL59UD-Q4_K_XL52IQ3_XXS39Q437Q8_035MXFP4284-bit28