Sep 18, 2026
Use cases
agentictool-usecoding
Summary
User reports DeepSeek V4.1 Flash at 300 t/s prefill and 16 t/s decode on an M3 Ultra 512GB, running the Q2 quant on the DS4 engine.
Tool calls were functional but the model made questionable choices, ignoring the code execution tool and searching the local drive for a remote file before finding the right path. The 18k-token Hermes prompt was insufficient to steer it.
At around 110k tokens of context the decode rate only just began to dip below 16 t/s, which the user notes is a key advantage of DS4 over llama.cpp. The GPU ran at about 95% throughout. The user compares it unfavorably to GLM 5.3 Flash, which puts out roughly 21-19 t/s, and is hesitant to try Q4.