llamaperf

Kimi K3

Moonshot AI · 1 report

Thin page (1 of 3 reports needed for indexing). Add yours.

Kimi K3

16× DGX Spark · vLLM

reported speed:
20.0 tokens/s generation · 750.0 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Full Kimi K3 model running on 16x DGX Spark cluster. Average 20+ tps, peak 38 tps, prefill 750 tps. First run, plans to optimize and publish vLLM image.