llamaperf

Qwen3.8 27B Swift-1.5

on NVIDIA RTX 5090 · NInfer · 262,144 ctx

Sep 25, 2026
Throughput
160.8 t/s gen · 3269.0 t/s pp
Quant
NVFP4 (NVFP4)
KV cache
k8v4
VRAM reported
32 GB

Use cases

codingagenticlong-contextvision

Summary

User reports Swift-1.5 Qwen3.8-27B at 160.8 t/s decode on a single RTX 5090 at 262,144 context. Setup is the NInfer v3 engine with an all-NVFP4 (W4A4 gs16) artifact plus a z-lab DFlash2 drafter at K=7, k8v4 KV cache, 18.0 GiB of weights in VRAM, 450 W power cap, and concurrency 4. Prefill at 200k context measured 3,269 t/s. IFBench prompt-strict 69.0, prompt-loose 72.7, instr-strict 70.4, instr-loose 73.6; GSM8K-200 95.0% (190/200); long-context needle at 250,031 tokens exact across 3 depths.