llamaperf

Qwen3.6 35B (3B active) Uncensored-HauhauCS-Aggressive

on NVIDIA RTX 4070 · ik_llama.cpp · 131,072 ctx

Tone: negative
Oct 4, 2026
Throughput
28.7 t/s gen
Quant
Q3_K_P (GGUF)
KV cache
q8_0
System RAM
32 GB

Summary

User reports Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive at 28.74 t/s generation on an RTX 4070 with a Ryzen 7 7700X and 32GB DDR5, roughly half the speed they get from llama.cpp. Setup is ik_llama.cpp with a Q3_K_P GGUF, q8_0 KV cache, 131072 context, and 12 threads; the same command under llama.cpp gave 63.82 t/s generation and 70.52 t/s prompt eval. The user is asking what could be done about the shortfall.