llamaperf

Qwen3.8 125B (6B active) Swift-1.5

on 2× NVIDIA RTX 3090 Ti · Strata · 393,216 ctx

Tone: positive
Oct 5, 2026
Throughput
105.0 t/s gen · 2000.0 t/s pp
Quant
IQ3-XXS (GGUF)
KV cache
int8
System RAM
64 GB

Summary

User reports Qwen3.8 Flash-Next Swift-1.5 at 105 t/s decode and around 2000 t/s prefill on an RTX 3090 Ti plus a V100 16GB. Setup is Strata 0.1.39 with IQ3-XXS weights, int8 KV cache, 393216 max context, yarn rope scaling at 1.5, and speculative decoding with --spec 4. Layers are split 2/3 on the 3090 Ti and 1/3 on the V100. User says the same model previously reached only around 50 t/s decode with llama.cpp layer splitting.