ds4
An inference engine for running open-weight LLMs locally.
11 community reports
This engine doesn't yet have an editorial profile on llamaperf. The community reports below show how it's been used in practice across different hardware.
Top GPUs running ds4
| GPU | VRAM | Reports | Median t/s, DeepSeek V4.1 Flash 552B · 16B active 2-bit |
|---|---|---|---|
| M3 Ultra 512GBapple | 512GB | 3 | 16.0 |
| M5 Pro 64GBapple | 64GB | 2 | no plain run of this model |
| AMD Strix Halo 128GBamd | 128GB | 1 | no plain run of this model |
| M1 Max 64GBapple | 64GB | 1 | no plain run of this model |
| M3 Max 96GBapple | 96GB | 1 | no plain run of this model |
| M3 Ultra 96GBapple | 96GB | 1 | no plain run of this model |
| M5 Max 128GBapple | 128GB | 1 | no plain run of this model |
| M5 Max 64GBapple | 64GB | 1 | no plain run of this model |
ds4 against other engines
Pairs of plain runs on the same card, of the same model size at the same bit class: one device, one request, no speculative decoding, the whole model in memory. Context length and build still differ between the two sides, and each side shows its own.
No matched pair yet. No card has plain runs of one model size at one bit class on ds4 and on another engine, so llamaperf can't say how it compares on speed. Add a run.
ds4 results by GPU
Every card people have run ds4 on, with each report's model, quant and speed, newest first. Runs on several cards, with speculative decoding, with batched requests or with part of the model in system RAM say so, since each describes a different setup.
ds4 on M3 Ultra 512GB3 reports
- DeepSeek V4.1 Flash 552B · 16B active · Q216.0 t/s, 300 t/s prompt
- DeepSeek V4.1 Flash 552B · 16B active · Q431.3 t/s, 813 t/s prompt(speculative)
- DeepSeek V4 Flash 284B · 13B active475 t/s prompt
ds4 on M5 Pro 64GB2 reports
- DeepSeek V4 Flash 284B · 13B active · Q8_015.0 t/s(part in system RAM)
- DeepSeek V4 Flash 284B · 13B active · mixed 4/8-bit2.0 t/s(part in system RAM)
ds4 on AMD Strix Halo 128GB1 report
- DeepSeek V4.1 Flash 552B · 16B active · Q27.1 t/s(part in system RAM)
ds4 on M1 Max 64GB1 report
- Qwen3.8 125B · 6B active · Q2_044.0 t/s, 328 t/s prompt(speculative)
ds4 on M3 Max 96GB1 report
ds4 on M3 Ultra 96GB1 report
- Qwen3.8 125B · 6B active · Q4667 t/s prompt(part in system RAM)
ds4 on M5 Max 128GB1 report
- DeepSeek V4.1 Flash 552B · 16B active · Q224.4 t/s, 636 t/s prompt(part in system RAM)
ds4 on M5 Max 64GB1 report
- DeepSeek V4 Flash 284B · 13B active · Q2-Q4 mixed imatrix30.0 t/s(part in system RAM)
ds4 on NVIDIA RTX 3080 20GB1 report
- DeepSeek V4 Flash 284B · 13B active · IQ2_M11.5 t/s, 300 t/s prompt(2 cards, part in system RAM)
ds4 on NVIDIA RTX 5060 8GB1 report
- DeepSeek V4.1 Flash 552B · 16B active · FP82.4 t/s(part in system RAM)
Top models on ds4
Frequently asked
Is ds4 faster than other engines?
llamaperf has no matched comparison for ds4 yet: no card has plain runs of the same model size at the same bit class on ds4 and on another engine. Speed claims about engines need that pairing, so this page doesn't make one.