Qwen3.8 125B (6B active) Flash-Next
on NVIDIA Jetson AGX Thor 128GB · llama.cpp · 65,536 ctx
Sep 27, 2026
Use cases
visioncodingagenticlong-context
Summary
User reports Qwen3.8-Flash-Next running on a Jetson AGX Thor 128GB with vision support, decoding at 15.3-18.5 tok/s on free-form text.
Setup is llama.cpp (qwen4exp branch, commit d4a943f plus cherry-pick 24ea62d and canreuse-v2.patch) with UD-Q4_K_XL GGUF, 65536 context, ngram-mod speculative decoding, and the 51B n-gram table offloaded to CPU/NVMe via -ot per_layer_token_embd=CPU -lm mmap, leaving about 80GB resident.
Prefill is 90-170 tok/s depending on caching; ngram-mod speculation reached 82.9 tok/s on verbatim code reproduction (91% draft acceptance, mean accepted span 59 tokens) but only fires on long untouched spans. The 120W power mode costs about 10% versus MAXN, and images cost a one-off 1-2s to encode without affecting decode speed.