llamaperf

Kimi K2.6

Moonshot AI · 2 reports

Thin page (2 of 3 reports needed for indexing). Add yours.

Kimi K2.6

RTX 5090 · llama.cpp

throughput:
471.4 t/s pp
quant:
IQ3_M (gguf)

LLM prompt processing benchmark with Kimi K2.5 IQ3_M (80GB offload) at 500W. RTX 5090 achieved 471.40 t/s PP. Also tested GLM 5.1 IQ4_NL (70GB offload) at 574.98 t/s PP. Comparison with RTX 6000 PRO MaxQ shunt modded.

throughput:
9.7 t/s gen · 264.0 t/s pp

32x AMD MI50 32GB GPUs. Prompt processing 264 t/s, generation 9.7 t/s. Model: Kimi K2.6.