Qwen3.8 27B
RTX 4080 · ExLlamaV3 · 131,072 ctx
- reported speed:
- 56.5 tokens/s generation · 980.2 tokens/s prompt processing
- quant:
- SC_3.00bpw_H4_V4 (EXL3)
- kv:
- 6,5
- mtp (multi-token prediction):
- on
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3.8 27B at 56.48 t/s decode and 980.2 t/s prefill on an RTX 4080 16GB at 131,072 context with 102,400 active input tokens. Setup is ExLlamaV3 1.5.0 with TabbyAPI, the turboderp SC_3.00bpw_H4_V4 EXL3 quant, a 6,5 KV cache, and MTP k=2 with a Q6 draft cache. VRAM peaked around 15.2GB of 16.4GB. Without MTP decode was 33.68 t/s, so MTP gave about a 68% decode increase at roughly a 5% prefill cost. Q4 draft cache was about 9% slower than Q6, and dynamic drafting was slower than fixed k=2.