Qwen3.6 35B (3B active) Uncensored-HauhauCS-Aggressive
on NVIDIA RTX 4070 · ik_llama.cpp · 131,072 ctx
Oct 4, 2026
Summary
User reports Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive at 28.74 t/s generation on an RTX 4070 with a Ryzen 7 7700X and 32GB DDR5, roughly half the speed they get from llama.cpp.
Setup is ik_llama.cpp with a Q3_K_P GGUF, q8_0 KV cache, 131072 context, and 12 threads; the same command under llama.cpp gave 63.82 t/s generation and 70.52 t/s prompt eval.
The user is asking what could be done about the shortfall.