llamaperf

Qwen3

Alibaba · 4 reports

By engine

EngineAvg t/sRangeN
text-generation-webui7.58–81
Ollama43.043–431
llama.cpp159.0159–1591
Tone: positive
reported speed:
43.0 tokens/s generation
quant:
Q4 (GGUF)
kv:
q8_0
rating:
5/5

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

visioncoding

Best overall VLM for OCR and detail extraction. Correctly read mixed-script text (Chinese + Latin) and caught fine details other models missed. Verbose output (1.4-2.2k tokens). Recommended as default for coding-assistant MCP.

Qwen3 235B (22B active)

RTX 3090 · text-generation-webui · 16,384 ctx

Tone: positive
reported speed:
7.5 tokens/s generation
quant:
Q4_K_M (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User praises Qwen3-235B-A22B for uncensored ERP roleplay, notes it's their daily driver for months. Mentions Step 3.7 Flash as alternative but censored and verbose. Also mentions Gemma 4 and MiniMax-M2.5 but not run. Setup: single RTX 3090 (24GB) plus 128GB system RAM (DDR4, dual Xeon), textgen with batch size 1024, 16384 context, autofit layers, threads 38/48. Reports ~75k pp and ~7.5 t/s generation (dips to ~7 t/s at 10k context).

Qwen3 8B

RTX 5090 · llama.cpp · 8,192 ctx

Tone: positive
reported speed:
159.0 tokens/s generation
quant:
BF16 (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

summarizationmultilingual

Benchmark of DSpark PC Tree speculative decoding fork. Best config PCTree k3/n16 achieved 159.00 tok/s vs plain 94.27 tok/s. Also tested Qwen3.8 27B Q4 with worse results.

Tone: positive
reported speed:
52.0 tokens/s generation
quant:
float8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Custom CUDA/C++ engine, 50-54 tok/s, 50% improvement over llama.cpp (33-34 tok/s).