Qwen3.8 27B
on NVIDIA V100 32GB · NInfer
Sep 11, 2026
Summary
User reports best-case MTP decode of 218 tok/s with 99.2% acceptance on a Volta V100.
Setup is the custom NInfer inference engine with software NVFP4.
Throughput is about 200 tok/s without context-copy speculation.