llamaperf

Qwen3.6 27B

on NVIDIA RTX 5090 Laptop 24GB · llama.cpp · 200,000 ctx

Tone: positive
Apr 28, 2026
Quant
IQ4_XS (GGUF)
KV cache
Q8
Rating
5/5
System RAM
64 GB
VRAM reported
24 GB

Use cases

codingtool-use

Summary

User reports Qwen 3.6 27B is excellent for pyspark/python and data transformation debugging, running on an ASUS ROG Strix SCAR 18 with an RTX 5090 laptop (24 GB VRAM) and 64 GB DDR5 RAM. Setup is llama.cpp with the IQ4_XS quant at 200k context and a Q8_0 KV cache. The user initially tried q4_k_m at q4_0. No tokens/sec is reported. The user is cancelling cloud subscriptions due to local performance.