llamaperf

Kimi K2.5

Moonshot AI · 1 report

Thin page (1 of 3 reports needed for indexing). Add yours.

Kimi K2.5

RTX 5090 · llama.cpp

reported speed:
471.4 tokens/s prompt processing
quant:
IQ3_M (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

LLM prompt processing benchmark with Kimi K2.5 IQ3_M (80GB offload) at 500W. RTX 5090 achieved 471.40 t/s PP. Also tested GLM 5.1 IQ4_NL (70GB offload) at 574.98 t/s PP. Comparison with RTX 6000 PRO MaxQ shunt modded.