DeepSeek V4 Flash 9B distill
M5 Pro 48GB · Ollama · 16,384 ctx
- reported speed:
- 44.0 tokens/s generation
- quant:
- Q4_K_M (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User compares DeepSeek V4 Flash 9B distill against its base model Qwen3.5 9B on a MacBook Pro M5 Pro with 48 GB RAM. Both models gave the same answers on 6/8 tasks, but the distill used fewer tokens (5480 vs 8975). Throughput was near identical (44 vs 41 tok/s). The distill was slower on open-ended tasks (1467 tokens vs 867 for a buffer overflow explanation). The base model failed arithmetic due to a token cap.