Tencent-HY3 295B (21B active)
8× RTX Pro 6000 Blackwell · llama.cpp · 262,144 ctx
- reported speed:
- 57.4 tokens/s generation
- quant:
- Q6_K (gguf)
- kv:
- F16
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
Also benchmarked Q4_K_M (67.3 t/s), Q3_K_L (63.3 t/s), IQ2_M (78.7 t/s). Second model: Nemotron-Labs-Audex-30B-A3B (30B MoE, ~3B active) with Q8_0 287 t/s, Q5_K_M 334 t/s, Q4_K_M 345 t/s, MXFP4_MOE 329 t/s on 2x RTX PRO 6000 Max-Q.