Qwen3-Next 80B (3B active)
on NVIDIA RTX 3090 · llama.cpp
Oct 6, 2026
Summary
User reports Qwen3-Next-80B-A3B at 94.5 tok/s decode on an RTX 3090 with 16 GB RAM, using a patched llama.cpp that substitutes missing experts instead of waiting for SSD reads.
Setup is llama.cpp with Q4_K_M, 1/4 of experts in VRAM, rest read from NVMe at ~5.7 GB/s, 16 threads, about 15.7 GB VRAM used.
Stock llama.cpp gave 31.8 tok/s in the same 16 GB case; the patch also reached 108.4 tok/s with plenty of RAM and 89 tok/s reading every miss from SSD. Perplexity was 1.6% higher than stock, GSM8K lost 1.8 points, and greedy generation ran 64-74 tok/s. The 16 GB case was simulated by locking RAM on a bigger machine.