llamaperf

Qwen3.6 27B

on NVIDIA RTX 3090 · ik_llama.cpp · 156,000 ctx

Tone: positive
May 19, 2026
Throughput
72.9 t/s gen · 1261.0 t/s pp
Quant
IQ4_KS (GGUF)
KV cache
Q8
Flash Attention
on
MTP (Multi-Token Prediction)
on
VRAM reported
24 GB

Use cases

coding

Summary

User reports Qwen3.6-27B-MTP-IQ4_KS.gguf at 72.9 t/s decode and 1261 tok/s prefill on an RTX 3090 24GB at 156k context. Setup is ik_llama.cpp with a q8_0/q8_0 KV cache, MTP, and vision on CPU. The user also tested llama.cpp and BeeLlama.