Unknown family
on NVIDIA RTX 5060 Ti 16GB · llama.cpp · 98,304 ctx
Sep 28, 2026
Summary
User reports failing to reproduce a claimed 50+ t/s on an RTX 5060 Ti 16GB, getting 20-35 t/s instead.
Setup is a freshly built llama.cpp fork (beellama.cpp) for sm120 with MTP speculative decoding, kvarn3 KV cache, 98304 context, and parallel 1.
The user questions whether the original high-throughput post is fake or if they are missing something.