Qwen3.6 27B
on NVIDIA RTX 3090 · ik_llama.cpp · 200,000 ctx
Sep 27, 2026
Use cases
coding
Summary
User benchmarks Qwen3.6 27B on a single RTX 3090 at an enforced 370 W power cap, comparing ik_llama.cpp with IQ4_KS against mainline llama.cpp with Q4_K_M. ik_llama.cpp reaches ~72 t/s code decode and ~60 t/s narrative decode versus ~58 and ~50 t/s for mainline, an 18-20% gain, with quality tied at 100-103/150 on an 8-pack and ~0.5 GB lower VRAM.
Setup uses q4_0 KV cache, MTP n=2, -np 1, and a 200K context configuration that fills to ~183K tokens. A two-stage ngram plus MTP speculator on the same engine lifts code decode to ~98 t/s at the same power cap, with narrative flat at ~59 t/s.
The user notes the earlier reported tie was a measurement artifact from a stale mainline container answering on the shared port, and that both engines are strongly power-sensitive between 230 W and 370 W.