DeepSeek V4.1 Flash 552B (16B active)
on M3 Ultra 512GB · ds4 · 300,000 ctx
Sep 15, 2026
Use cases
agenticcoding
Summary
User reports DeepSeek V4.1 Flash at 31.3 t/s decode at 8k context on an M3 Ultra 512GB, up from 16.6 t/s upstream.
Setup is a forked ds4 engine with Q4 GGUF weights, DSpark multi-token speculative decoding, and a disk KV cache. At 300k context decode is 28.3 t/s, prefill on a 62k prompt is 813 t/s, and TTFT on a 23k system prompt is 31.1 s.
With DSpark on, code generation reaches 40.5 t/s and agent answer phases 41.3 t/s. A 91-minute agent turn decoded 101k tokens with 4.6M tokens prefilled at 99.5% cache hit and 56 tool calls. Output is byte-identical to upstream under greedy decode.