DeepSeek V4 Flash 284B (13B active)
A100 40GB · llama.cpp
- reported speed:
- 16.1 tokens/s generation
- quant:
- Q8_K_XL (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports all experts on CPU with only 15.8 GB of 40 GB VRAM used.