llamaperf

Qwen3-Coder

1 report

As of 7 Oct 2026, Qwen3-Coder 30B · 3B active at 4-bit on the hardware it is most run on, with the median of plain runs (one device, one request, no speculative decoding, the whole model in its memory):

Qwen3-Coder VRAM requirements by size and quant →

How does Qwen3-Coder run on your hardware?

Pick your GPU or Mac, then read the speeds people reported for Qwen3-Coder on it, or estimate memory fit and speed.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Run Qwen3-Coder yourself? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance for Qwen3-Coder

Filter this model’s reports by setup →
Thin page (1 of 3 reports needed for indexing). Add yours.
reported speed:
66.1 tokens/s generation
quant:
Q4_K_M (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User benchmarks Qwen3 Coder 30B A3B at 66.11 t/s generation on one AMD Radeon Instinct Mi50 32GB. Setup is llama.cpp build 128d522c (6686) with the ROCm backend, Q4_K_M quant (17.28 GiB), -ngl 99, 112 threads and --numa distribute on a dual Xeon 8480+ host. Other quants on the same card: Q5_0 65.15 t/s, Q6_K 62.49 t/s, Q8_0 64.20 t/s. BF16 (56.89 GiB) required two GPUs and ran 41.41 t/s. With 16K tokens of output the rate falls to about 20 t/s.

Sep 30, 2026

Get a weekly email of new Qwen3-Coder reports on any GPU.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Qwen3-Coder 30B (3B active)
AMD MI50 32GB
Q4_K_M
llama.cpp
Not reported66.1 tokens/s