llamaperf

Qwen3.8 27B

on NVIDIA V100 32GB · NInfer

Tone: positive
Sep 11, 2026
Throughput
218.0 t/s gen
Quant
NVFP4
MTP (Multi-Token Prediction)
on
VRAM reported
32 GB

Summary

User reports best-case MTP decode of 218 tok/s with 99.2% acceptance on a Volta V100. Setup is the custom NInfer inference engine with software NVFP4. Throughput is about 200 tok/s without context-copy speculation.