DeepSeek V4 Flash 284B (13B active)
RTX A6000 48GB · 1,000,000 ctx
- reported speed:
- 17.2 tokens/s generation · 70.0 tokens/s prompt processing
- quant:
- Q8_K_XL
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports prompt processing in the high 70s t/s, dropping to the mid 30s at 300k context. The full 1M context fits in 48 GB of VRAM, but prompt processing is slow.