- reported speed:
- 11.0 tokens/s generation
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports DeepSeek-V4-Flash (284B) at about 11 t/s on a 64GB MacBook. The model is 165GB on disk and does not fit in RAM, so it runs via SSD streaming with a 2-bit file. Quality matches official perplexity (6.1250 vs 6.1262). Higher quality settings give 1-2 t/s. An 8GB cache gave 2.04 t/s versus 1.23 t/s with 32GB. Speculative decoding and prefetching did not help.