llamaperf

Qwen3.8 27B

on NVIDIA RTX 5090 Laptop 24GB · llama.cpp · 65,536 ctx

Tone: mixed
Oct 7, 2026
Throughput
30.0 t/s gen
Quant
Q4_K_M (GGUF)
KV cache
8bit
VRAM reported
24 GB

Use cases

coding

Summary

User reports Qwen3.8 27B at a steady 30 t/s on an RTX 5090 Laptop 24GB. Setup is llama.cpp with Q4_K_M and 8-bit KV cache at 65k context, using 21 GB of VRAM. The user also tested Unsloth UD_Q5_K_XL at 27 t/s, and NInfer models reaching 82.54 t/s (8-bit, 132k context), 92.4 t/s (NVFP4, 32k context), and 86.63 t/s (NVFP4, 4-bit, 64k context). The user says the laptop is too slow for coding and recommends a DGX Spark instead.