Qwopus3.6 27B Coder-Compat
NVIDIA V100 32GB · llama.cpp · 8,192 ctx
- reported speed:
- 40.3 tokens/s generation · 527.0 tokens/s prompt processing
- quant:
- Q4_K_M (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User benchmarks Qwopus3.6 27B Coder-Compat at 40.3 t/s decode on a Tesla V100 32GB, with 527 t/s prompt throughput. Setup is llama.cpp v9836 (upstream-dflash branch) with the Q4_K_M GGUF at 8K context, flash attention on, and MTP n=3 speculative decoding; two decode runs averaged. Baseline without speculative decoding was 31.2 t/s decode. On an RTX 3060 12GB, Qwopus3.6 35B-A3B Coder-MTP ran 25.5 t/s baseline but dropped to 16.9 t/s with MTP n=3, a regression.