Qwen3.8 27B
RTX 2080 Ti · NInfer · 128,000 ctx
- reported speed:
- 45.0 tokens/s generation
- quant:
- W8A16
- kv:
- Q8
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
Speculative decoding with MTP3 draft window yields ~456 tok/s, ~65% acceptance rate. Standard autoregressive ~25 tok/s. VRAM usage ~17.5 GiB with draft weights, leaving ~4.5-5.0 GiB for KV cache.