llamaperf

Qwen3-Next 80B (3B active)

on Unknown GPU · llama.cpp · 32,768 ctx

Tone: positive
Oct 3, 2026
Throughput
19.1 t/s gen · 130.0 t/s pp
Quant
Q4_K_M (GGUF)
System RAM
256 GB

Use cases

coding

Summary

User reports Qwen3-Next-80B-A3B at 130 t/s prefill and 19.1 t/s generation on a CPU-only Lenovo ThinkStation P5 with Intel Xeon w5-2555X and 256 GB DDR5 ECC. Setup is llama.cpp built with GGML_NATIVE=ON, Q4_K_M quant, 32k context, mlocked into RAM, no GPU offload. The user also benchmarks Qwen3-Coder-30B-A3B at 166 t/s prefill and 33.7 t/s generation, gpt-oss-120b at 98 t/s prefill and 19.5 t/s generation, and Qwen3-235B-A22B at 24.8 t/s prefill and 5.6 t/s generation. An NVIDIA T1000 8GB was tested and found to slow prefill versus CPU-only, so it was removed from inference. The user notes that -ub tuning varies per model and that Q8 quant of Qwen3-Next-80B scored 59/61 on a coding suite versus 55-57 for Q4_K_M.