- reported speed:
- 9.7 tokens/s generation · 264.0 tokens/s prompt processing
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Kimi K2.6 at 264 t/s prompt processing and 9.7 t/s generation on 32x AMD MI50 32GB GPUs.