llamaperf

StepFun 3.7

StepFun · 1 report

Thin page (1 of 3 reports needed for indexing). Add yours.

StepFun 3.7

M5 Max 128GB · llama.cpp · 65,536 ctx

reported speed:
33.9 tokens/s generation
quant:
Q4_K_S (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Benchmark at context length 65536. Prompt processing t/s at various context lengths: 128 (62.80), 2048 (60.52), 8192 (57.32), 16384 (51.71), 32768 (45.43), 65536 (33.92). Memory peak ~120+ GB, system sluggish but usable.