Qwen3.8 125B (6B active) Flash-Next
on NVIDIA Jetson AGX Thor 128GB · TensorFold · 262,144 ctx
Oct 4, 2026
Summary
User reports Qwen3.8-Flash-Next NVFP4 at 36.7 tok/s one-user decode and 899 tok/s prefill at 32k on a Jetson AGX Thor 128GB.
Setup is TensorFold 0.6.0 with int8 KV cache, full 262,144-token window, --parallel 4. Prefill measured at 926 / 899 / 798 tok/s for 8k / 32k / 128k tokens; combined decode at 1 / 2 / 3 / 4 users is 34.8 / 53.7 / 70.3 / 76.7 tok/s; first answer to a 60k-token prompt takes 69.7 s.
User compares against vLLM's FP8 build on the same box (38.4 tok/s decode, 2,101 tok/s prefill at 32k) and describes a local Thor prompt path with gated runs r09 and r11 reaching 1,549 tok/s prefill at 32k. 128 GB and 273 GB/s are published Thor specs, not re-measured.