llamaperf

Qwen3.8 27B Swift

on NVIDIA RTX 5090 · NInfer

Tone: positive
Sep 23, 2026

Use cases

agentictool-use

Summary

User reports a native ninfer implementation of a Jev-like decision API running Qwen3.8 27B (Swift fine-tune) on an RTX 5090, scoring 84.4% (195/231) on the JevBench v1.2 public set with three in-flight decisions and a 19 s wall clock. Setup drives the native POST /v1/decisions route directly rather than the chat-completions path; latency p50 is 127 ms and p95 is 699 ms, with hard-tier p95 prefill-dominated by ~3.7k-token states and tail latency inflated by a single serialized decision worker. Calibration is ECE 0.045 and Brier 0.081; the user calls it a dirty POC with bugs and incorrect allocation still to fix, and notes latency is not great but accuracy is almost on par.