llamaperf

NVIDIA RTX 5090 Laptop 24GB

NVIDIA · 24GB · 4 reports

As of 7 Oct 2026, the models most run on the NVIDIA RTX 5090 Laptop 24GB, with the median of plain runs (one device, one request, no speculative decoding):

Engines people use on it: llama.cpp 1

Run models on your NVIDIA RTX 5090 Laptop 24GB? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the NVIDIA RTX 5090 Laptop 24GB

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 24 GB of VRAM.

Qwen3.8 27B

NVIDIA RTX 5090 Laptop 24GB · llama.cpp · 65,536 ctx

Tone: mixed
reported speed:
30.0 tokens/s generation
quant:
Q4_K_M (GGUF)
kv:
8bit

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen3.8 27B at a steady 30 t/s on an RTX 5090 Laptop 24GB. Setup is llama.cpp with Q4_K_M and 8-bit KV cache at 65k context, using 21 GB of VRAM. The user also tested Unsloth UD_Q5_K_XL at 27 t/s, and NInfer models reaching 82.54 t/s (8-bit, 132k context), 92.4 t/s (NVFP4, 32k context), and 86.63 t/s (NVFP4, 4-bit, 64k context). The user says the laptop is too slow for coding and recommends a DGX Spark instead.

Oct 7, 2026
Tone: mixed

User reports a DIY Jev-like inference setup using unmodified open weight LLMs, evaluated on a 32,235-example benchmark. Qwen3.6 35B-A3B reached 75.5% accuracy at ~5.3 req/s on a laptop RTX 5090 24GB, with Qwen3 27B at 75.3% ~2.9 req/s and Qwen3-4B at 65.0% ~27 req/s. The approach uses boolean verification of candidate answers via true/false logits, batched through llama.cpp, with no NLI fine-tuning or classifier head. User notes the benchmark is not perfectly apples-to-apples and expresses skepticism about Jev hype.

Sep 20, 2026
Tone: mixed
reported speed:
3.6 tokens/s generation
quant:
INT4_SYM (OpenVINO IR)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3-14B Heretic at 3.56 t/s on an Intel AI Boost NPU37XX in an Acer Predator Helios 16 AI laptop with 32 GB RAM and an RTX 5090 Laptop 24GB. Setup is OpenVINO GenAI with the machine-made-Fibre INT4_SYM OpenVINO IR export, group_size=-1, all_layers=true, loaded directly via openvino_genai.LLMPipeline on NPU with no conversion. Load took 33.52 s and 96 output tokens took 26.98 s; peak RAM was about 15.6 GB with a 9.7 GB working set. User also benchmarked Qwen3-8B INT4 SYM on the same NPU at 6.72 t/s with 6.15 s load, and found Qwen3-30B-A3B MoE impractical due to host memory pressure during OpenVINO preparation.

Sep 18, 2026

Qwen3.6 27B

NVIDIA RTX 5090 Laptop 24GB · llama.cpp · 200,000 ctx

Tone: positive
quant:
IQ4_XS (GGUF)
kv:
Q8
rating:
5/5
codingtool-use

User reports Qwen 3.6 27B is excellent for pyspark/python and data transformation debugging, running on an ASUS ROG Strix SCAR 18 with an RTX 5090 laptop (24 GB VRAM) and 64 GB DDR5 RAM. Setup is llama.cpp with the IQ4_XS quant at 200k context and a Q8_0 KV cache. The user initially tried q4_k_m at q4_0. No tokens/sec is reported. The user is cancelling cloud subscriptions due to local performance.

Apr 28, 2026

Get a weekly email of new NVIDIA RTX 5090 Laptop 24GB reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Qwen3.8 27B
NVIDIA RTX 5090 Laptop 24GB
Q4_K_M
llama.cpp
65,53630.0 tokens/s
Qwen3 14B Heretic
NVIDIA RTX 5090 Laptop 24GB
INT4_SYM
OpenVINO GenAI
Not reported3.6 tokens/s