Qwen3-Next 80B (3B active)
NVIDIA RTX 3090 · llama.cpp
- reported speed:
- 94.5 tokens/s generation
- quant:
- Q4_K_M (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3-Next-80B-A3B at 94.5 tok/s decode on an RTX 3090 with 16 GB RAM, using a patched llama.cpp that substitutes missing experts instead of waiting for SSD reads. Setup is llama.cpp with Q4_K_M, 1/4 of experts in VRAM, rest read from NVMe at ~5.7 GB/s, 16 threads, about 15.7 GB VRAM used. Stock llama.cpp gave 31.8 tok/s in the same 16 GB case; the patch also reached 108.4 tok/s with plenty of RAM and 89 tok/s reading every miss from SSD. Perplexity was 1.6% higher than stock, GSM8K lost 1.8 points, and greedy generation ran 64-74 tok/s. The 16 GB case was simulated by locking RAM on a bigger machine.