llamaperf
Oct 5, 2026
Throughput
2.4 t/s gen
Quant
FP8 (safetensors)
System RAM
31 GB
VRAM reported
8 GB

Summary

User reports DeepSeek-V4.1-Flash (552B backbone + 196B Engram, 384 experts top-6, FP8 + FP4) running on a single RTX 5060 8GB at ~1.6 tokens/s from disk only and ~2.4 tokens/s with a 16 GB RAM cache. Setup streams experts and Engram rows from NVMe using DeepSeek's reference code with a modified storage layer; only the dense part lives on the GPU (cap --vram_gb 7.3), with 240 experts (4.2 GiB) read per token. Batch size 1, text only, context limit 8192, prefill max 700 tokens. Live chat via the OpenAI-compatible server with a 16 GB RAM cache reaches 2.43 tokens/s; peak VRAM 7.06 GiB. A 311-token prompt takes 17.9 s. The user notes the speed-ups changed nothing in output (136/136 tokens identical) and that DeepSeek's untouched reference could not be run side by side because it cannot load the model in 8 GB.