Qwen3.6 27B
4× RTX 5060 Ti 16GB · 256,000 ctx
- throughput:
- 52.2 t/s gen · 608.0 t/s pp
- quant:
- Q8
- kv:
- F16
- mtp (multi-token prediction):
- on
Benchmark on Vast AI instance with 4x RTX 5060 Ti 16GB. Q8 quant, FP16 KV cache, MTP enabled. 256K context. Cold prefill 608 t/s, decode 52.2 t/s. User considers this excellent for $2K hardware.