llamaperf

Local model performance on your hardware

Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.

What runs on your hardware?

Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Local model performance reports from the community

These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →

Model: Qwen3-Next
Compare setup details

Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.

Tone: positive
reported speed:
94.5 tokens/s generation
quant:
Q4_K_M (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3-Next-80B-A3B at 94.5 tok/s decode on an RTX 3090 with 16 GB RAM, using a patched llama.cpp that substitutes missing experts instead of waiting for SSD reads. Setup is llama.cpp with Q4_K_M, 1/4 of experts in VRAM, rest read from NVMe at ~5.7 GB/s, 16 threads, about 15.7 GB VRAM used. Stock llama.cpp gave 31.8 tok/s in the same 16 GB case; the patch also reached 108.4 tok/s with plenty of RAM and 89 tok/s reading every miss from SSD. Perplexity was 1.6% higher than stock, GSM8K lost 1.8 points, and greedy generation ran 64-74 tok/s. The 16 GB case was simulated by locking RAM on a bigger machine.

Oct 6, 2026

Qwen3-Next 80B (3B active)

Unknown GPU · llama.cpp · 32,768 ctx

Tone: positive
reported speed:
19.1 tokens/s generation · 130.0 tokens/s prompt processing
quant:
Q4_K_M (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen3-Next-80B-A3B at 130 t/s prefill and 19.1 t/s generation on a CPU-only Lenovo ThinkStation P5 with Intel Xeon w5-2555X and 256 GB DDR5 ECC. Setup is llama.cpp built with GGML_NATIVE=ON, Q4_K_M quant, 32k context, mlocked into RAM, no GPU offload. The user also benchmarks Qwen3-Coder-30B-A3B at 166 t/s prefill and 33.7 t/s generation, gpt-oss-120b at 98 t/s prefill and 19.5 t/s generation, and Qwen3-235B-A22B at 24.8 t/s prefill and 5.6 t/s generation. An NVIDIA T1000 8GB was tested and found to slow prefill versus CPU-only, so it was removed from inference. The user notes that -ub tuning varies per model and that Q8 quant of Qwen3-Next-80B scored 59/61 on a coding suite versus 55-57 for Q4_K_M.

Oct 3, 2026

Qwen3-Next Flash

6× BC-250 · llama.cpp · 100,000 ctx

Tone: positive
reported speed:
28.0 tokens/s generation
quant:
IQ2_XS (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3-Next Flash IQ2_XS at around 28 tok/s for short generation on a cluster of 4 BC-250 boards, dropping to 24 tok/s at 50k context with around 115 t/s prefill. Setup is llama.cpp with Vulkan and RPC over 1Gb Ethernet, 100k context, across 6 BC-250 ex-mining boards in an ASRock 4U12G case. The other two boards run Qwen3.6 35B Q4 at 60 tok/s with 100k context and 450 t/s prefill.

Oct 2, 2026
reported speed:
116.0 tokens/s generation · 6800-7950 tokens/s prompt processing
quant:
W4A16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3-Next-80B at 116 t/s single-stream decode on a single NVIDIA CMP 170HX with 64 GB HBM2e, at a 150 W cap. Setup is vLLM 0.27.1 with W4A16 weights (40.9 GB), torch 2.13.0+cu130, CUDA 13.0, Ubuntu 26.04 LTS, on an AMD Ryzen Threadripper PRO 3945WX with 128 GB DDR4 ECC. Prefill over about 8.9k tokens measured 6800 to 7950 t/s; power draw 137 to 145 W. Aggregate throughput at 8 concurrent requests was 352 t/s, saturating at 4 slots. Also measured on the same card for context: Ornith-1.5-35B FP8 at 122.5 t/s and Qwen3.8-27B W4A16 with DFlash2 at 127 t/s single-stream.

Sep 23, 2026
Showing 1–4 of 4
Page 1 of 1

Community benchmarks snapshot

Records by GPU

NVIDIA RTX 3090165NVIDIA RTX 5090114AMD Strix Halo 128GB82NVIDIA DGX Spark57NVIDIA RTX 5060 Ti 16GB53NVIDIA RTX Pro 6000 Blackwell51NVIDIA RTX 3060 12GB46AMD Radeon AI PRO R9700 32GB43NVIDIA RTX 409037NVIDIA RTX 5070 Ti30

Records by model

1397 total
Qwen3.8781
Qwen3.6170
DeepSeek V4 Flash121
Gemma 461
Qwen3.529
Qwen322
other213

Records by engine

1074 total
llama.cpp570
vLLM153
Strata47
NInfer40
Ollama34
other230

Use cases

coding 453agentic 287long-context 208tool-use 120vision 85summarization 45math 36creative-writing 30multilingual 19text-generation 9rp 6reasoning 3
coding453agentic287long-context208tool-use120vision85summarization45math36creative-writing30

Median t/s by GPU

On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking

RTX 509093RTX Pro 600067M5 Ultra 256GB50RX 7900 XTX41RTX 309036V100 32GB33RTX 409032RX 7800 XT 16GB30RTX 5090 Laptop 24GB30Radeon AI PRO R9700 32GB29

Reports by model size

Qwen3.8 27B468Qwen3.8 125B · 6B active283DeepSeek V4 Flash 284B · 13B active102Qwen3.6 35B · 3B active100Qwen3.6 27B68Gemma 4 26B · 4B active25DeepSeek V4.1 Flash 552B · 16B active22GLM-5.3 320B · 18B active18

Quants

Q4_K_M119NVFP495IQ4_XS57Q4_K_XL56UD-Q4_K_XL49IQ3_XXS38Q437Q8_0354-bit24Q6_K23