llamaperf

Qwen3.8 27B

on NVIDIA RTX Pro 4000 Blackwell · NInfer · 128,000 ctx

Tone: positive
Sep 7, 2026
Throughput
67.0 t/s gen · 785.0 t/s pp
KV cache
Q8
VRAM reported
24 GB

Summary

User reports 67 t/s with MTP3 speculative decoding enabled, against 24.4 tok/s without MTP. Setup uses an INT8 KV cache with group-64. The figure comes from a 128K NIAH benchmark with 130,048 prompt tokens, where MTP acceptance was 100% on a deterministic answer. Roughly 727 MiB of VRAM was left.