- reported speed:
- 1.0 tokens/s generation · 50.0 tokens/s prompt processing
- quant:
- 4bit
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports DeepSeek V4 Flash 0731 at about 50 t/s prefill and about 1 t/s decode on an M5 Air with 32 GB. Setup uses the streamed experts trick with a model of roughly 300B total parameters.