llamaperf

Qwen3-Next

Alibaba · 4 reports

As of 7 Oct 2026, Qwen3-Next 80B · 3B active at 4-bit on the hardware it is most run on, with the median of plain runs (one device, one request, no speculative decoding, the whole model in its memory):

Qwen3-Next VRAM requirements by size and quant →

How does Qwen3-Next run on your hardware?

Pick your GPU or Mac, then read the speeds people reported for Qwen3-Next on it, or estimate memory fit and speed.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Run Qwen3-Next yourself? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance for Qwen3-Next

Filter this model’s reports by setup →
Tone: positive
reported speed:
94.5 tokens/s generation
quant:
Q4_K_M (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3-Next-80B-A3B at 94.5 tok/s decode on an RTX 3090 with 16 GB RAM, using a patched llama.cpp that substitutes missing experts instead of waiting for SSD reads. Setup is llama.cpp with Q4_K_M, 1/4 of experts in VRAM, rest read from NVMe at ~5.7 GB/s, 16 threads, about 15.7 GB VRAM used. Stock llama.cpp gave 31.8 tok/s in the same 16 GB case; the patch also reached 108.4 tok/s with plenty of RAM and 89 tok/s reading every miss from SSD. Perplexity was 1.6% higher than stock, GSM8K lost 1.8 points, and greedy generation ran 64-74 tok/s. The 16 GB case was simulated by locking RAM on a bigger machine.

Oct 6, 2026

Qwen3-Next 80B (3B active)

Unknown GPU · llama.cpp · 32,768 ctx

Tone: positive
reported speed:
19.1 tokens/s generation · 130.0 tokens/s prompt processing
quant:
Q4_K_M (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen3-Next-80B-A3B at 130 t/s prefill and 19.1 t/s generation on a CPU-only Lenovo ThinkStation P5 with Intel Xeon w5-2555X and 256 GB DDR5 ECC. Setup is llama.cpp built with GGML_NATIVE=ON, Q4_K_M quant, 32k context, mlocked into RAM, no GPU offload. The user also benchmarks Qwen3-Coder-30B-A3B at 166 t/s prefill and 33.7 t/s generation, gpt-oss-120b at 98 t/s prefill and 19.5 t/s generation, and Qwen3-235B-A22B at 24.8 t/s prefill and 5.6 t/s generation. An NVIDIA T1000 8GB was tested and found to slow prefill versus CPU-only, so it was removed from inference. The user notes that -ub tuning varies per model and that Q8 quant of Qwen3-Next-80B scored 59/61 on a coding suite versus 55-57 for Q4_K_M.

Oct 3, 2026

Qwen3-Next Flash

6× BC-250 · llama.cpp · 100,000 ctx

Tone: positive
reported speed:
28.0 tokens/s generation
quant:
IQ2_XS (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3-Next Flash IQ2_XS at around 28 tok/s for short generation on a cluster of 4 BC-250 boards, dropping to 24 tok/s at 50k context with around 115 t/s prefill. Setup is llama.cpp with Vulkan and RPC over 1Gb Ethernet, 100k context, across 6 BC-250 ex-mining boards in an ASRock 4U12G case. The other two boards run Qwen3.6 35B Q4 at 60 tok/s with 100k context and 450 t/s prefill.

Oct 2, 2026
reported speed:
116.0 tokens/s generation · 6800-7950 tokens/s prompt processing
quant:
W4A16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3-Next-80B at 116 t/s single-stream decode on a single NVIDIA CMP 170HX with 64 GB HBM2e, at a 150 W cap. Setup is vLLM 0.27.1 with W4A16 weights (40.9 GB), torch 2.13.0+cu130, CUDA 13.0, Ubuntu 26.04 LTS, on an AMD Ryzen Threadripper PRO 3945WX with 128 GB DDR4 ECC. Prefill over about 8.9k tokens measured 6800 to 7950 t/s; power draw 137 to 145 W. Aggregate throughput at 8 concurrent requests was 352 t/s, saturating at 4 slots. Also measured on the same card for context: Ornith-1.5-35B FP8 at 122.5 t/s and Qwen3.8-27B W4A16 with DFlash2 at 127 t/s single-stream.

Sep 23, 2026
Engines people run Qwen3-Next with
EngineReports
llama.cpp3
vLLM1

Get a weekly email of new Qwen3-Next reports on any GPU.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Qwen3-Next 80B (3B active)
NVIDIA RTX 3090
Q4_K_M
llama.cpp
Not reported94.5 tokens/s
Qwen3-Next 80B (3B active)
NVIDIA CMP 170HX 64GB (unlocked)
W4A16
vLLM
Not reported116.0 tokens/s