Qwen3.6 27B
on NVIDIA RTX 3090 · ik_llama.cpp · 156,000 ctx
May 19, 2026
Use cases
coding
Summary
User reports Qwen3.6-27B-MTP-IQ4_KS.gguf at 72.9 t/s decode and 1261 tok/s prefill on an RTX 3090 24GB at 156k context.
Setup is ik_llama.cpp with a q8_0/q8_0 KV cache, MTP, and vision on CPU.
The user also tested llama.cpp and BeeLlama.