llamaperf

Local model performance on your hardware

Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.

What runs on your hardware?

Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Local model performance reports from the community

These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →

GPU: NVIDIA Jetson AGX Thor 128GB
Compare setup details

Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.

reported speed:
36.7 tokens/s generation · 899.0 tokens/s prompt processing
quant:
NVFP4 (NVFP4)
kv:
int8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8-Flash-Next NVFP4 at 36.7 tok/s one-user decode and 899 tok/s prefill at 32k on a Jetson AGX Thor 128GB. Setup is TensorFold 0.6.0 with int8 KV cache, full 262,144-token window, --parallel 4. Prefill measured at 926 / 899 / 798 tok/s for 8k / 32k / 128k tokens; combined decode at 1 / 2 / 3 / 4 users is 34.8 / 53.7 / 70.3 / 76.7 tok/s; first answer to a 60k-token prompt takes 69.7 s. User compares against vLLM's FP8 build on the same box (38.4 tok/s decode, 2,101 tok/s prefill at 32k) and describes a local Thor prompt path with gated runs r09 and r11 reaching 1,549 tok/s prefill at 32k. 128 GB and 273 GB/s are published Thor specs, not re-measured.

Oct 4, 2026
reported speed:
25.0 tokens/s generation · 3557.0 tokens/s prompt processing
quant:
NVFP4 (NVFP4)
kv:
8-bit

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks Qwen3.8-27B at 24.96 t/s generation and 3557 t/s prompt processing on a Jetson AGX Thor 128GB, using the Mjolnir vLLM image with the FA4 GEMV decode kernel. Setup is vLLM 0.30.0 with NVFP4 weights and an 8-bit KV cache at 8K context, single request. The GEMV kernel is default-on and dispatches only on M=1, head_dim=256, GQA shapes. The GEMV leg wins at c=1 but regresses at c=4 (57.74 t/s at 8K context, −15.4% versus stock vLLM), which the user attributes to an open investigation. Stock vLLM reaches 24.43 t/s and 2468 t/s prompt at the same setting.

Oct 4, 2026
Tone: positive
reported speed:
15.3-18.5 tokens/s generation · 90-170 tokens/s prompt processing
quant:
Q4_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

visioncodingagenticlong-context

User reports Qwen3.8-Flash-Next running on a Jetson AGX Thor 128GB with vision support, decoding at 15.3-18.5 tok/s on free-form text. Setup is llama.cpp (qwen4exp branch, commit d4a943f plus cherry-pick 24ea62d and canreuse-v2.patch) with UD-Q4_K_XL GGUF, 65536 context, ngram-mod speculative decoding, and the 51B n-gram table offloaded to CPU/NVMe via -ot per_layer_token_embd=CPU -lm mmap, leaving about 80GB resident. Prefill is 90-170 tok/s depending on caching; ngram-mod speculation reached 82.9 tok/s on verbatim code reproduction (91% draft acceptance, mean accepted span 59 tokens) but only fires on long untouched spans. The 120W power mode costs about 10% versus MAXN, and images cost a one-off 1-2s to encode without affecting decode speed.

Sep 27, 2026
reported speed:
139.1 tokens/s generation
quant:
NVFP4 (NVFP4)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.6-35B-A3B-NVFP4 at 139.1 tok/s on a single NVIDIA Jetson AGX Thor with 117 GB unified memory. Setup is vLLM built from source for sm_110a with DFlash speculative decoding (12 tokens), marlin MoE backend, flash_attn attention backend, 65536 context length, and 0.78 GPU memory utilization. The post also benchmarks Qwen3.5-4B-NVFP4 at 155.8 tok/s, Qwen3.6-27B-NVFP4 at 50.1 tok/s, and Qwen3.5-122B-A10B-NVFP4 at 52.6 tok/s, all at concurrency 1. The 122B requires cutlass MoE and TRITON_ATTN due to a Marlin crash at 256 experts, and achieves 27-42 tok/s with DFlash versus 10.9 tok/s autoregressive.

Sep 27, 2026
Showing 1–4 of 4
Page 1 of 1

Community benchmarks snapshot

Records by GPU

NVIDIA RTX 3090165NVIDIA RTX 5090114AMD Strix Halo 128GB82NVIDIA DGX Spark57NVIDIA RTX 5060 Ti 16GB53NVIDIA RTX Pro 6000 Blackwell51NVIDIA RTX 3060 12GB46AMD Radeon AI PRO R9700 32GB43NVIDIA RTX 409037NVIDIA RTX 5070 Ti30

Records by model

1397 total
Qwen3.8781
Qwen3.6170
DeepSeek V4 Flash121
Gemma 461
Qwen3.529
Qwen322
other213

Records by engine

1074 total
llama.cpp570
vLLM153
Strata47
NInfer40
Ollama34
other230

Use cases

coding 453agentic 287long-context 208tool-use 120vision 85summarization 45math 36creative-writing 30multilingual 19text-generation 9rp 6reasoning 3
coding453agentic287long-context208tool-use120vision85summarization45math36creative-writing30

Median t/s by GPU

On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking

RTX 509093RTX Pro 600067M5 Ultra 256GB50RX 7900 XTX41RTX 309036V100 32GB33RTX 409032RX 7800 XT 16GB30RTX 5090 Laptop 24GB30Radeon AI PRO R9700 32GB29

Reports by model size

Qwen3.8 27B468Qwen3.8 125B · 6B active283DeepSeek V4 Flash 284B · 13B active102Qwen3.6 35B · 3B active100Qwen3.6 27B68Gemma 4 26B · 4B active25DeepSeek V4.1 Flash 552B · 16B active22GLM-5.3 320B · 18B active18

Quants

Q4_K_M119NVFP495IQ4_XS57Q4_K_XL56UD-Q4_K_XL49IQ3_XXS38Q437Q8_0354-bit24Q6_K23