llamaperf
Sep 27, 2026
Throughput
92.0 t/s gen
Quant
MXFP4 (MXFP4)

Summary

User reports Kimi K3 (2.8T parameters) at 92 tok/s decode on 8x B300 via Modal. Setup is vLLM with MXFP4 weights, tensor parallel 8, cold boot about 27 min for a 1.56 TB load. TTFT is 0.92 to 1.02 s and average decode over 4 prompts is 83 tok/s. Cost is $56.79 per hour, $190 per million output tokens, about $36 per run, or $1,363 a day left warm. User also ran Unsloth's Dynamic GGUF 1-bit UD-IQ1_S (594 GB) on 8x A100-80GB via llama.cpp at about 9 tok/s with TTFT 7 to 60 s, $19.99 per hour and about $620 per million tokens, 3.3x more expensive per token. Quality at 1-bit was fine.