Qwen3.8 125B (6B active) Flash-Next
on NVIDIA RTX 3080 Laptop 16GB · llama.cpp · 65,000 ctx
Sep 29, 2026
Summary
User reports running Qwen3.8-Flash-Next on an RTX 3080 Laptop GPU with 16GB VRAM at roughly 10-20 tok/s generation and an estimated 100-200 tok/s prompt processing, with about 65k context and only 1-2GB system RAM usage.
Setup uses a modified llama.cpp. The prompt processing figure is an estimate based on cloud GPU calculations, not a direct measurement on the laptop.
The user was mid-benchmark when their 180W power adapter cable failed and is asking for donations to replace it, promising to release the technique and source code regardless.