- reported speed:
- 11.1 tokens/s generation
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3.8 Flash-Next 176B at 11.09 tok/s decode on an RTX 3080 Laptop 16GB with 32GB system RAM and SSD. Setup is TensorSharp with MoE-aware unified scheduling across cache, VRAM, system RAM, and SSD; quant and context length not stated. User compares against Strata, which reached 10.24 tok/s decode and 62.15s whole-process time versus TensorSharp's 16.54s; TensorSharp GPU peak was 14,832.5 MiB and OS peak working set 19.74 GiB.