Qwen3.6 35B (3B active)
2× NVIDIA P102-100 · llama.cpp · 32,768 ctx
- reported speed:
- 23.5 tokens/s generation · 432.3 tokens/s prompt processing
- quant:
- IQ4_XS (gguf)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports 70 t/s total across 3 concurrent users (23.3 t/s each) with 32K context per user. Uses two P102-100 cards (10GB each) for $100 total. Prompt processing speed 432 t/s. Model is Qwen3.6-35B-A3B at IQ4_XS quantization.