llamaperf

Qwen3.8 27B

on 2× NVIDIA RTX 5070 Ti · NInfer · 7,680 ctx

Tone: positive
Sep 28, 2026
Throughput
220.0 t/s gen · 4956.0 t/s pp
Quant
NVFP4 (NVFP4)

Use cases

codingagentic

Summary

User reports Qwen3.8-27B at 220.0 tok/s decode with MTP3 structured output on 2x RTX 5070 Ti, matching an RTX 5090's 219.8 tok/s on the same official NVFP4 weights. Setup is the NInfer tensor-parallel fork v0.2.3 with NVFP4 weights, 7,680-token context, C=1, uncapped clocks, no P2P over PCIe 5.0 x8/x8. Without speculation decode is 70.9 tok/s and prefill 4,956 tok/s (59% of the 5090's 8,340). Against llama.cpp on the same two cards: cold prefill 4,577 vs 2,197 tok/s, decode at 184k context 122.9 vs 64.0 tok/s, and a 16-turn agent session to ~94k context 46.5 s vs 82.8 s. QUASAR-QAT weights reach 254.5 tok/s decode but are lighter (17.0 vs 20.9 GiB).