Qwen3.8 27B
on NVIDIA RTX Pro 4000 Blackwell · NInfer · 128,000 ctx
Sep 7, 2026
Summary
User reports 67 t/s with MTP3 speculative decoding enabled, against 24.4 tok/s without MTP.
Setup uses an INT8 KV cache with group-64.
The figure comes from a 128K NIAH benchmark with 130,048 prompt tokens, where MTP acceptance was 100% on a deterministic answer. Roughly 727 MiB of VRAM was left.