Qwen3.5 9B
2× D700 12GB · llama.cpp · 70,000 ctx
- reported speed:
- 11.0 tokens/s generation
- quant:
- Q4 (gguf)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
coding
Also tested Qwen 2.5 coder q4 at 22 t/s. User compares Qwen 3.5 favorably to Claude Sonnet 4.6 for planning tasks.