llamaperf
Sep 29, 2026
Throughput
10-20 t/s gen · 100-200 t/s pp
VRAM reported
16 GB

Summary

User reports running Qwen3.8-Flash-Next on an RTX 3080 Laptop GPU with 16GB VRAM at roughly 10-20 tok/s generation and an estimated 100-200 tok/s prompt processing, with about 65k context and only 1-2GB system RAM usage. Setup uses a modified llama.cpp. The prompt processing figure is an estimate based on cloud GPU calculations, not a direct measurement on the laptop. The user was mid-benchmark when their 180W power adapter cable failed and is asking for donations to replace it, promising to release the technique and source code regardless.