llamaperf

DeepSeek V4.1 Flash 552B (16B active)

on M3 Ultra 512GB · ds4 · 300,000 ctx

Tone: positive
Sep 15, 2026
Throughput
31.3 t/s gen · 813.0 t/s pp
Quant
Q4 (GGUF)
MTP (Multi-Token Prediction)
on
System RAM
512 GB

Use cases

agenticcoding

Summary

User reports DeepSeek V4.1 Flash at 31.3 t/s decode at 8k context on an M3 Ultra 512GB, up from 16.6 t/s upstream. Setup is a forked ds4 engine with Q4 GGUF weights, DSpark multi-token speculative decoding, and a disk KV cache. At 300k context decode is 28.3 t/s, prefill on a 62k prompt is 813 t/s, and TTFT on a 23k system prompt is 31.1 s. With DSpark on, code generation reaches 40.5 t/s and agent answer phases 41.3 t/s. A 91-minute agent turn decoded 101k tokens with 4.6M tokens prefilled at 99.5% cache hit and 56 tool calls. Output is byte-identical to upstream under greedy decode.