Qwen3-Next 80B (3B active)
on Unknown GPU · llama.cpp · 32,768 ctx
Oct 3, 2026
Use cases
coding
Summary
User reports Qwen3-Next-80B-A3B at 130 t/s prefill and 19.1 t/s generation on a CPU-only Lenovo ThinkStation P5 with Intel Xeon w5-2555X and 256 GB DDR5 ECC.
Setup is llama.cpp built with GGML_NATIVE=ON, Q4_K_M quant, 32k context, mlocked into RAM, no GPU offload.
The user also benchmarks Qwen3-Coder-30B-A3B at 166 t/s prefill and 33.7 t/s generation, gpt-oss-120b at 98 t/s prefill and 19.5 t/s generation, and Qwen3-235B-A22B at 24.8 t/s prefill and 5.6 t/s generation. An NVIDIA T1000 8GB was tested and found to slow prefill versus CPU-only, so it was removed from inference. The user notes that -ub tuning varies per model and that Q8 quant of Qwen3-Next-80B scored 59/61 on a coding suite versus 55-57 for Q4_K_M.