llamaperf

Tencent-HY3

Tencent · 2 reports

Thin page (2 of 3 reports needed for indexing). Add yours.
reported speed:
57.4 tokens/s generation
quant:
Q6_K (gguf)
kv:
F16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Also benchmarked Q4_K_M (67.3 t/s), Q3_K_L (63.3 t/s), IQ2_M (78.7 t/s). Second model: Nemotron-Labs-Audex-30B-A3B (30B MoE, ~3B active) with Q8_0 287 t/s, Q5_K_M 334 t/s, Q4_K_M 345 t/s, MXFP4_MOE 329 t/s on 2x RTX PRO 6000 Max-Q.

Tone: positive
reported speed:
32.4 tokens/s generation
quant:
UD128 (gguf)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

tool-use

User reports ~32.4 t/s generation at empty context, 16.3 t/s at 16K context. With MTP n=2, peaks at 38 t/s. Prefill speeds: 528 t/s (pp512 empty), 124 t/s (pp512 at 16K). Model is Tencent HY3 295B-A21B MoE. Quant is UD128 (107GB). Uses llama.cpp with Metal, q8_0 KV cache, 24K context. User compares favorably to DeepSeek V4 Flash.