llamaperf

NVIDIA Jetson AGX Thor 128GB

NVIDIA · 128GB unified memory · 4 reports

As of 7 Oct 2026, the models most run on the NVIDIA Jetson AGX Thor 128GB, with the median of plain runs (one device, one request, no speculative decoding):

Engines people use on it: vLLM 2

Run models on your NVIDIA Jetson AGX Thor 128GB? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the NVIDIA Jetson AGX Thor 128GB

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 128 GB of unified memory.

reported speed:
36.7 tokens/s generation · 899.0 tokens/s prompt processing
quant:
NVFP4 (NVFP4)
kv:
int8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8-Flash-Next NVFP4 at 36.7 tok/s one-user decode and 899 tok/s prefill at 32k on a Jetson AGX Thor 128GB. Setup is TensorFold 0.6.0 with int8 KV cache, full 262,144-token window, --parallel 4. Prefill measured at 926 / 899 / 798 tok/s for 8k / 32k / 128k tokens; combined decode at 1 / 2 / 3 / 4 users is 34.8 / 53.7 / 70.3 / 76.7 tok/s; first answer to a 60k-token prompt takes 69.7 s. User compares against vLLM's FP8 build on the same box (38.4 tok/s decode, 2,101 tok/s prefill at 32k) and describes a local Thor prompt path with gated runs r09 and r11 reaching 1,549 tok/s prefill at 32k. 128 GB and 273 GB/s are published Thor specs, not re-measured.

Oct 4, 2026
reported speed:
25.0 tokens/s generation · 3557.0 tokens/s prompt processing
quant:
NVFP4 (NVFP4)
kv:
8-bit

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks Qwen3.8-27B at 24.96 t/s generation and 3557 t/s prompt processing on a Jetson AGX Thor 128GB, using the Mjolnir vLLM image with the FA4 GEMV decode kernel. Setup is vLLM 0.30.0 with NVFP4 weights and an 8-bit KV cache at 8K context, single request. The GEMV kernel is default-on and dispatches only on M=1, head_dim=256, GQA shapes. The GEMV leg wins at c=1 but regresses at c=4 (57.74 t/s at 8K context, −15.4% versus stock vLLM), which the user attributes to an open investigation. Stock vLLM reaches 24.43 t/s and 2468 t/s prompt at the same setting.

Oct 4, 2026
Tone: positive
reported speed:
15.3-18.5 tokens/s generation · 90-170 tokens/s prompt processing
quant:
Q4_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

visioncodingagenticlong-context

User reports Qwen3.8-Flash-Next running on a Jetson AGX Thor 128GB with vision support, decoding at 15.3-18.5 tok/s on free-form text. Setup is llama.cpp (qwen4exp branch, commit d4a943f plus cherry-pick 24ea62d and canreuse-v2.patch) with UD-Q4_K_XL GGUF, 65536 context, ngram-mod speculative decoding, and the 51B n-gram table offloaded to CPU/NVMe via -ot per_layer_token_embd=CPU -lm mmap, leaving about 80GB resident. Prefill is 90-170 tok/s depending on caching; ngram-mod speculation reached 82.9 tok/s on verbatim code reproduction (91% draft acceptance, mean accepted span 59 tokens) but only fires on long untouched spans. The 120W power mode costs about 10% versus MAXN, and images cost a one-off 1-2s to encode without affecting decode speed.

Sep 27, 2026
reported speed:
139.1 tokens/s generation
quant:
NVFP4 (NVFP4)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.6-35B-A3B-NVFP4 at 139.1 tok/s on a single NVIDIA Jetson AGX Thor with 117 GB unified memory. Setup is vLLM built from source for sm_110a with DFlash speculative decoding (12 tokens), marlin MoE backend, flash_attn attention backend, 65536 context length, and 0.78 GPU memory utilization. The post also benchmarks Qwen3.5-4B-NVFP4 at 155.8 tok/s, Qwen3.6-27B-NVFP4 at 50.1 tok/s, and Qwen3.5-122B-A10B-NVFP4 at 52.6 tok/s, all at concurrency 1. The 122B requires cutlass MoE and TRITON_ATTN due to a Marlin crash at 256 experts, and achieves 27-42 tok/s with DFlash versus 10.9 tok/s autoregressive.

Sep 27, 2026

Get a weekly email of new NVIDIA Jetson AGX Thor 128GB reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Qwen3.8 125B (6B active) Flash-Next
NVIDIA Jetson AGX Thor 128GB
NVFP4
TensorFold
262,14436.7 tokens/s
Qwen3.8 27B
NVIDIA Jetson AGX Thor 128GB
NVFP4
vLLM
8,19225.0 tokens/s
Qwen3.6 35B (3B active)
NVIDIA Jetson AGX Thor 128GB
NVFP4
vLLM
65,536139.1 tokens/s