DeepSeek V4.1 Flash 552B (16B active)
on M5 Max 128GB · ds4 · 2,048 ctx
Oct 7, 2026
Summary
User reports DeepSeek V4.1 Flash at Q2 running on an M5 Max 128GB Mac, with a fork of ds4 reaching 24.4 t/s generation at 2048 context versus 12.1 t/s for stock ds4.
Setup is ds4 with a 341 GB GGUF (152 GB weights plus 189 GB Engram tables) streamed from SSD, no KV quant, bit-exact output. The fork doubles steady decode to 25.7 t/s at 2048 context and 22.0 t/s at 32768 context, and prefill from 16k to 32k context rises from 404 t/s to 636 t/s.
Adding a byte-identical external Thunderbolt 5 SSD copy cuts time to first token on a 3.5K prompt from 13.2 s to 11.2 s and after an 8K context from 1.38 s to 0.18 s; decode is unchanged. The enclosure runs at PCIe 4.0 x4, about half the internal SSD speed. A two-SSD attempt in upstream ds4's mmap path was bit-exact but 13-38% slower on prefill and was not submitted.