llamaperf

Qwopus3.6

1 report

Qwopus3.6 VRAM requirements by size and quant →

How does Qwopus3.6 run on your hardware?

Pick your GPU or Mac, then read the speeds people reported for Qwopus3.6 on it, or estimate memory fit and speed.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Run Qwopus3.6 yourself? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance for Qwopus3.6

Filter this model’s reports by setup →
Thin page (1 of 3 reports needed for indexing). Add yours.
reported speed:
40.3 tokens/s generation · 527.0 tokens/s prompt processing
quant:
Q4_K_M (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagentic

User benchmarks Qwopus3.6 27B Coder-Compat at 40.3 t/s decode on a Tesla V100 32GB, with 527 t/s prompt throughput. Setup is llama.cpp v9836 (upstream-dflash branch) with the Q4_K_M GGUF at 8K context, flash attention on, and MTP n=3 speculative decoding; two decode runs averaged. Baseline without speculative decoding was 31.2 t/s decode. On an RTX 3060 12GB, Qwopus3.6 35B-A3B Coder-MTP ran 25.5 t/s baseline but dropped to 16.9 t/s with MTP n=3, a regression.

Sep 23, 2026

Get a weekly email of new Qwopus3.6 reports on any GPU.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Qwopus3.6 27B Coder-Compat
NVIDIA V100 32GB
Q4_K_M
llama.cpp
8,19240.3 tokens/s