Qwen3.6 27B
2× RTX 3090 Ti · llama.cpp · 196,608 ctx
- reported speed:
- 100.0 tokens/s generation
- quant:
- Q8_0 (gguf)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports tensor split-mode raising throughput from 70+ t/s to 100+ t/s, with a peak of 130 t/s. Power draw is 750W+.