Qwen3.6 35B (3B active)
NVIDIA RTX 2080 Ti 22GB (modded) · llama.cpp · 16,384 ctx
- reported speed:
- 69.0 tokens/s generation
- quant:
- UD-IQ4_XS (GGUF)
- kv:
- q8_0
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3.6-35B-A3B at 4-bit (UD-IQ4_XS) scoring 89.63% pass@1 on HumanEval at 69 tok/s on a single RTX 2080 Ti 22GB. Setup is llama.cpp with UD-IQ4_XS dynamic 4-bit (4.25 bpw), q8_0 KV cache, 16384 context, flash-attn on, whole model resident in VRAM with no offload. A routing patch (MoE expansion, 20 experts instead of 8 on layers 25-39) reached 90.85% pass@1 at 56 tok/s, a 19% decode speed cost; the user calls the +2 problems within statistical noise and notes the tests are original HumanEval, not EvalPlus.