- reported speed:
- 20.0 tokens/s generation · 750.0 tokens/s prompt processing
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
Full Kimi K3 model running on 16x DGX Spark cluster. Average 20+ tps, peak 38 tps, prefill 750 tps. First run, plans to optimize and publish vLLM image.