Qwen3.6 27B
2× RTX 3090 Ti · llama.cpp · 196,608 ctx
- throughput:
- 100.0 t/s gen
- quant:
- Q8_0 (gguf)
Tensor split-mode improved t/s from 70+ to 100+. Peak 130 t/s reported. Power draw 750W+.
NVIDIA · 24GB · 1 report
2× RTX 3090 Ti · llama.cpp · 196,608 ctx
Tensor split-mode improved t/s from 70+ to 100+. Peak 130 t/s reported. Power draw 750W+.