llamaperf
Oct 3, 2026
Throughput
11.1 t/s gen
System RAM
32 GB
VRAM reported
16 GB

Summary

User reports Qwen3.8 Flash-Next 176B at 11.09 tok/s decode on an RTX 3080 Laptop 16GB with 32GB system RAM and SSD. Setup is TensorSharp with MoE-aware unified scheduling across cache, VRAM, system RAM, and SSD; quant and context length not stated. User compares against Strata, which reached 10.24 tok/s decode and 62.15s whole-process time versus TensorSharp's 16.54s; TensorSharp GPU peak was 14,832.5 MiB and OS peak working set 19.74 GiB.