Qwen3.6 27B
on NVIDIA RTX 5090 Laptop 24GB · llama.cpp · 200,000 ctx
Apr 28, 2026
Use cases
codingtool-use
Summary
User reports Qwen 3.6 27B is excellent for pyspark/python and data transformation debugging, running on an ASUS ROG Strix SCAR 18 with an RTX 5090 laptop (24 GB VRAM) and 64 GB DDR5 RAM.
Setup is llama.cpp with the IQ4_XS quant at 200k context and a Q8_0 KV cache. The user initially tried q4_k_m at q4_0.
No tokens/sec is reported. The user is cancelling cloud subscriptions due to local performance.