llamaperf

Qwen3.8 27B Swift-Genesis

on 2× NVIDIA RTX 5060 Ti 16GB · llama.cpp · 262,144 ctx

Tone: mixed
Sep 27, 2026
Throughput
76.0 t/s gen
MTP (Multi-Token Prediction)
on
VRAM reported
32 GB

Use cases

visionagenticlong-context

Summary

User reports Qwen3.8 27B Swift-Genesis at 76 t/s in their benchmark on 2x RTX 5060 Ti 16GB, with a range of 40-113 t/s and 45-65 t/s under normal agentic work. Setup is llama.cpp with GGUF weights at 262k context, MTP4 speculative decoding, and vision offloaded to system RAM. The user notes the model fits under 17GB with vision, allowing full 262k context on 32GB VRAM, and asks what the tradeoff of this model is compared to other Swift NVFP4 models.