Qwen3.8 125B (6B active) Flash-Next
on NVIDIA RTX 3080 Laptop 16GB · TensorSharp
Oct 3, 2026
Summary
User reports Qwen3.8 Flash-Next 176B at 11.09 tok/s decode on an RTX 3080 Laptop 16GB with 32GB system RAM and SSD.
Setup is TensorSharp with MoE-aware unified scheduling across cache, VRAM, system RAM, and SSD; quant and context length not stated.
User compares against Strata, which reached 10.24 tok/s decode and 62.15s whole-process time versus TensorSharp's 16.54s; TensorSharp GPU peak was 14,832.5 MiB and OS peak working set 19.74 GiB.