Qwen3.6 35B (3B active)
NVIDIA RTX 4060 Ti 8GB · llama.cpp · 131,072 ctx
- reported speed:
- 52-65 tokens/s generation
- quant:
- Q4_K_XL (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3.6-35B-A3B at 52-65 tok/s on an RTX 4060 Ti 8GB with 64 GB system RAM. Setup is llama.cpp with Q4_K_XL at 131k context, experts offloaded to system RAM and the rest in VRAM, on headless Linux. The user compares download defaults (~25 tok/s), tuned Windows (39-45 tok/s), and tuned headless Linux (52-65 tok/s). They also report Qwen3.8-Flash-Next 125B at 17-19 tok/s and Ternary Bonsai 27B at 36 tok/s.