DeepSeek V4 Flash 284B (13B active)
M1 Max 64GB · llama.cpp · 65,536 ctx
- reported speed:
- 8.0 tokens/s generation · 30.0 tokens/s prompt processing
- quant:
- IQ3-XXS
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports DeepSeek-V4-Flash-0731 at about 8 t/s decode and about 30 t/s prefill on an M1 Max 64GB. Setup is a patched llama.cpp with the IQ3-XXS quant at 104 GB and context limited to 64k.