Llama 3.1 405B
Unknown GPU
- reported speed:
- 1.2 tokens/s generation
- quant:
- Q4_K_M (gguf)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Llama 405B Q4_K_M at 1.2 t/s two years ago, and contrasts it with current speeds of 30-100 t/s for newer models including Kimi K2.6, DeepSeek V4 Flash, MiniMax 2.7, Step 3.5 Flash and Qwen3.5-397B. User also mentions running Qwen3.6-36B at 50 t/s for a few hundred dollars.