DeepSeek V4.1 Flash
M2 Max 64GB · Whallm
- reported speed:
- 2.0 tokens/s generation
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports DeepSeek V4.1 Flash at 1.8-2.2 t/s on an M2 Max 64GB. Setup is the Whallm app with MLX, loading parts of the model from SSD, around 33 GiB peak MLX memory, with 1K-16K input tokens. The app also adds built-in throughput benchmarks and an OpenAI-compatible API.