DeepSeek V4 Flash
on 2× NVIDIA RTX 3090 · FreeToken
Sep 27, 2026
Summary
User reports DeepSeek V4 Flash at 5.58 tok/s on 2x RTX 3090 with the offload MoE backend, versus 0.67 tok/s with the auto-selected hybrid backend.
Setup is FreeToken 0.1.2 with fp8 dense weights and fp4 experts, TP=2, moe-cache-auto resolving to 624 slots (5.7% residency). Decode measured over a streamed 60-token completion, first token excluded, after a warmup request.
The auto-selected hybrid backend is 8.3x slower than offload. Raising memory-ratio to 0.95 grew the expert pool to 779 slots (7.1% residency) for only +6.6% throughput (5.95 tok/s). Doubling CPU threads from 23 to 44 yielded only +15%. The user notes TP=2 is a regression for offloaded MoE per issue #62, so absolute numbers are pessimistic.