- reported speed:
- 43.0 tokens/s generation
- quant:
- Q4 (GGUF)
- kv:
- q8_0
- rating:
- 5/5
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
visioncoding
Best overall VLM for OCR and detail extraction. Correctly read mixed-script text (Chinese + Latin) and caught fine details other models missed. Verbose output (1.4-2.2k tokens). Recommended as default for coding-assistant MCP.
- reported speed:
- 7.5 tokens/s generation
- quant:
- Q4_K_M (gguf)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User praises Qwen3-235B-A22B for uncensored ERP roleplay, notes it's their daily driver for months. Mentions Step 3.7 Flash as alternative but censored and verbose. Also mentions Gemma 4 and MiniMax-M2.5 but not run. Setup: single RTX 3090 (24GB) plus 128GB system RAM (DDR4, dual Xeon), textgen with batch size 1024, 16384 context, autofit layers, threads 38/48. Reports ~75k pp and ~7.5 t/s generation (dips to ~7 t/s at 10k context).
- reported speed:
- 159.0 tokens/s generation
- quant:
- BF16 (gguf)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
summarizationmultilingual
Benchmark of DSpark PC Tree speculative decoding fork. Best config PCTree k3/n16 achieved 159.00 tok/s vs plain 94.27 tok/s. Also tested Qwen3.8 27B Q4 with worse results.
- reported speed:
- 52.0 tokens/s generation
- quant:
- float8
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
Custom CUDA/C++ engine, 50-54 tok/s, 50% improvement over llama.cpp (33-34 tok/s).