Qwen3.8 27B
on 2× NVIDIA RTX 5070 Ti · NInfer · 7,680 ctx
Sep 28, 2026
Use cases
codingagentic
Summary
User reports Qwen3.8-27B at 220.0 tok/s decode with MTP3 structured output on 2x RTX 5070 Ti, matching an RTX 5090's 219.8 tok/s on the same official NVFP4 weights.
Setup is the NInfer tensor-parallel fork v0.2.3 with NVFP4 weights, 7,680-token context, C=1, uncapped clocks, no P2P over PCIe 5.0 x8/x8. Without speculation decode is 70.9 tok/s and prefill 4,956 tok/s (59% of the 5090's 8,340).
Against llama.cpp on the same two cards: cold prefill 4,577 vs 2,197 tok/s, decode at 184k context 122.9 vs 64.0 tok/s, and a 16-turn agent session to ~94k context 46.5 s vs 82.8 s. QUASAR-QAT weights reach 254.5 tok/s decode but are lighter (17.0 vs 20.9 GiB).