llamaperf

Step-3.5-Flash 196B

on 3× NVIDIA RTX 3090 · llama.cpp · 16,384 ctx

Tone: mixed
Oct 8, 2026
Throughput
17.5 t/s gen
Quant
IQ4_XS (GGUF)
KV cache
Q8_0
System RAM
64 GB

Summary

User reports Step-3.5-Flash at about 17/18 t/s on llama.cpp with an RTX 3090 and two RTX 5060 Ti 16GB cards, 64GB RAM and a Ryzen 9 5950X. Setup is llama.cpp with an IQ4_XS GGUF around 96GB, 16384 context and Q8_0 KV cache, with the model split across GPU and system RAM. The same model on ik_llama.cpp reaches only 10 t/s and is a bit unstable, so the user asks whether the multi-GPU mix is the problem or whether the setup can be refined.