llamaperf

NVIDIA RTX 3080 Laptop 16GB

NVIDIA · 16GB · 2 reports

Engines people use on it: TensorSharp 1

Run models on your NVIDIA RTX 3080 Laptop 16GB? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the NVIDIA RTX 3080 Laptop 16GB

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 16 GB of VRAM.

This page is thin (2 of 3 reports needed for indexing). Help fill it in.
Tone: positive
reported speed:
11.1 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 Flash-Next 176B at 11.09 tok/s decode on an RTX 3080 Laptop 16GB with 32GB system RAM and SSD. Setup is TensorSharp with MoE-aware unified scheduling across cache, VRAM, system RAM, and SSD; quant and context length not stated. User compares against Strata, which reached 10.24 tok/s decode and 62.15s whole-process time versus TensorSharp's 16.54s; TensorSharp GPU peak was 14,832.5 MiB and OS peak working set 19.74 GiB.

Oct 3, 2026
reported speed:
10-20 tokens/s generation · 100-200 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports running Qwen3.8-Flash-Next on an RTX 3080 Laptop GPU with 16GB VRAM at roughly 10-20 tok/s generation and an estimated 100-200 tok/s prompt processing, with about 65k context and only 1-2GB system RAM usage. Setup uses a modified llama.cpp. The prompt processing figure is an estimate based on cloud GPU calculations, not a direct measurement on the laptop. The user was mid-benchmark when their 180W power adapter cable failed and is asking for donations to replace it, promising to release the technique and source code regardless.

Sep 29, 2026

Get a weekly email of new NVIDIA RTX 3080 Laptop 16GB reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Qwen3.8 125B (6B active) Flash-Next
NVIDIA RTX 3080 Laptop 16GB
Not reported
TensorSharp
Not reported11.1 tokens/s