Step-3.5-Flash 196B
3× NVIDIA RTX 3090 · llama.cpp · 16,384 ctx
- reported speed:
- 17.5 tokens/s generation
- quant:
- IQ4_XS (GGUF)
- kv:
- Q8_0
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Step-3.5-Flash at about 17/18 t/s on llama.cpp with an RTX 3090 and two RTX 5060 Ti 16GB cards, 64GB RAM and a Ryzen 9 5950X. Setup is llama.cpp with an IQ4_XS GGUF around 96GB, 16384 context and Q8_0 KV cache, with the model split across GPU and system RAM. The same model on ik_llama.cpp reaches only 10 t/s and is a bit unstable, so the user asks whether the multi-GPU mix is the problem or whether the setup can be refined.