Qwen3.8 27B
NVIDIA RTX 5090 Laptop 24GB · llama.cpp · 65,536 ctx
- reported speed:
- 30.0 tokens/s generation
- quant:
- Q4_K_M (GGUF)
- kv:
- 8bit
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3.8 27B at a steady 30 t/s on an RTX 5090 Laptop 24GB. Setup is llama.cpp with Q4_K_M and 8-bit KV cache at 65k context, using 21 GB of VRAM. The user also tested Unsloth UD_Q5_K_XL at 27 t/s, and NInfer models reaching 82.54 t/s (8-bit, 132k context), 92.4 t/s (NVFP4, 32k context), and 86.63 t/s (NVFP4, 4-bit, 64k context). The user says the laptop is too slow for coding and recommends a DGX Spark instead.