Qwen3.8 27B Swift-Genesis
on 2× NVIDIA RTX 5060 Ti 16GB · llama.cpp · 262,144 ctx
Sep 27, 2026
Use cases
visionagenticlong-context
Summary
User reports Qwen3.8 27B Swift-Genesis at 76 t/s in their benchmark on 2x RTX 5060 Ti 16GB, with a range of 40-113 t/s and 45-65 t/s under normal agentic work.
Setup is llama.cpp with GGUF weights at 262k context, MTP4 speculative decoding, and vision offloaded to system RAM.
The user notes the model fits under 17GB with vision, allowing full 262k context on 32GB VRAM, and asks what the tradeoff of this model is compared to other Swift NVFP4 models.