llamaperf

Unknown family

on NVIDIA RTX 5060 Ti 16GB · llama.cpp · 98,304 ctx

Tone: mixed
Sep 28, 2026
Throughput
20-35 t/s gen
KV cache
kvarn3
Flash Attention
on

Summary

User reports failing to reproduce a claimed 50+ t/s on an RTX 5060 Ti 16GB, getting 20-35 t/s instead. Setup is a freshly built llama.cpp fork (beellama.cpp) for sm120 with MTP speculative decoding, kvarn3 KV cache, 98304 context, and parallel 1. The user questions whether the original high-throughput post is fake or if they are missing something.