Qwen3.8 27B
on NVIDIA RTX 5090 Laptop 24GB · llama.cpp · 65,536 ctx
Oct 7, 2026
Use cases
coding
Summary
User reports Qwen3.8 27B at a steady 30 t/s on an RTX 5090 Laptop 24GB.
Setup is llama.cpp with Q4_K_M and 8-bit KV cache at 65k context, using 21 GB of VRAM.
The user also tested Unsloth UD_Q5_K_XL at 27 t/s, and NInfer models reaching 82.54 t/s (8-bit, 132k context), 92.4 t/s (NVFP4, 32k context), and 86.63 t/s (NVFP4, 4-bit, 64k context). The user says the laptop is too slow for coding and recommends a DGX Spark instead.