llamaperf

Qwen3.8 125B (6B active) Flash-Next

on NVIDIA Jetson AGX Thor 128GB · llama.cpp · 65,536 ctx

Tone: positive
Sep 27, 2026
Throughput
15.3-18.5 t/s gen · 90-170 t/s pp
Quant
Q4_K_XL (GGUF)
System RAM
128 GB

Use cases

visioncodingagenticlong-context

Summary

User reports Qwen3.8-Flash-Next running on a Jetson AGX Thor 128GB with vision support, decoding at 15.3-18.5 tok/s on free-form text. Setup is llama.cpp (qwen4exp branch, commit d4a943f plus cherry-pick 24ea62d and canreuse-v2.patch) with UD-Q4_K_XL GGUF, 65536 context, ngram-mod speculative decoding, and the 51B n-gram table offloaded to CPU/NVMe via -ot per_layer_token_embd=CPU -lm mmap, leaving about 80GB resident. Prefill is 90-170 tok/s depending on caching; ngram-mod speculation reached 82.9 tok/s on verbatim code reproduction (91% draft acceptance, mean accepted span 59 tokens) but only fires on long untouched spans. The 120W power mode costs about 10% versus MAXN, and images cost a one-off 1-2s to encode without affecting decode speed.