- reported speed:
- 30.0 tokens/s generation
- quant:
- Q2-Q4 mixed imatrix (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports about 30 t/s on an M5 Max with a quantized DeepSeek V4 Flash running on the DS4 engine. The user states this is double the speed of llama.cpp. The user asks for feedback from CUDA and ROCm users.