Qwopus3.6 27B Coder-Compat
on NVIDIA V100 32GB · llama.cpp · 8,192 ctx
Sep 23, 2026
Use cases
codingagentic
Summary
User benchmarks Qwopus3.6 27B Coder-Compat at 40.3 t/s decode on a Tesla V100 32GB, with 527 t/s prompt throughput.
Setup is llama.cpp v9836 (upstream-dflash branch) with the Q4_K_M GGUF at 8K context, flash attention on, and MTP n=3 speculative decoding; two decode runs averaged.
Baseline without speculative decoding was 31.2 t/s decode. On an RTX 3060 12GB, Qwopus3.6 35B-A3B Coder-MTP ran 25.5 t/s baseline but dropped to 16.9 t/s with MTP n=3, a regression.