StepFun 3.7
M5 Max 128GB · llama.cpp · 65,536 ctx
- reported speed:
- 33.9 tokens/s generation
- quant:
- Q4_K_S (gguf)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
Benchmark at context length 65536. Prompt processing t/s at various context lengths: 128 (62.80), 2048 (60.52), 8192 (57.32), 16384 (51.71), 32768 (45.43), 65536 (33.92). Memory peak ~120+ GB, system sluggish but usable.