llamaperf

Step-3.5-Flash

1 report

Step-3.5-Flash VRAM requirements by size and quant →

How does Step-3.5-Flash run on your hardware?

Pick your GPU or Mac, then read the speeds people reported for Step-3.5-Flash on it, or estimate memory fit and speed.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Run Step-3.5-Flash yourself? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance for Step-3.5-Flash

Filter this model’s reports by setup →
Thin page (1 of 3 reports needed for indexing). Add yours.

Step-3.5-Flash 196B

3× NVIDIA RTX 3090 · llama.cpp · 16,384 ctx

Tone: mixed
reported speed:
17.5 tokens/s generation
quant:
IQ4_XS (GGUF)
kv:
Q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Step-3.5-Flash at about 17/18 t/s on llama.cpp with an RTX 3090 and two RTX 5060 Ti 16GB cards, 64GB RAM and a Ryzen 9 5950X. Setup is llama.cpp with an IQ4_XS GGUF around 96GB, 16384 context and Q8_0 KV cache, with the model split across GPU and system RAM. The same model on ik_llama.cpp reaches only 10 t/s and is a bit unstable, so the user asks whether the multi-GPU mix is the problem or whether the setup can be refined.

Oct 8, 2026

Get a weekly email of new Step-3.5-Flash reports on any GPU.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Step-3.5-Flash 196B
3× NVIDIA RTX 3090
IQ4_XS
llama.cpp
16,38417.5 tokens/s