llamaperf

Llama 3.1

Meta · 1 report

Llama 3.1 VRAM requirements by size and quant →
Thin page (1 of 3 reports needed for indexing). Add yours.

Llama 3.1 405B

Unknown GPU

Tone: positive
reported speed:
1.2 tokens/s generation
quant:
Q4_K_M (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Llama 405B Q4_K_M at 1.2 t/s two years ago, and contrasts it with current speeds of 30-100 t/s for newer models including Kimi K2.6, DeepSeek V4 Flash, MiniMax 2.7, Step 3.5 Flash and Qwen3.5-397B. User also mentions running Qwen3.6-36B at 50 t/s for a few hundred dollars.

May 4, 2026