llamaperf
Oct 3, 2026
Throughput
29.3 t/s gen
Quant
PTQ1_0 (GGUF)

Use cases

math

Summary

User reports Ternary Bonsai 2 27B at 29.27 t/s on an RTX 5060 Laptop 8 GB under Windows, using a Prism llama.cpp build. Setup is llama.cpp with PTQ1_0 (5.53 GiB), temp 0, seed 42, a 2048 reasoning budget and a 3072 token cap on 50 MATH-500 questions. A −2 logit bias on "wait", "maybe" and "perhaps" (nine token ids) dropped accuracy from 44/50 to 43/50 and raised average tokens from 845.2 to 872.3; the biased run measured 29.28 t/s. The user calls it one deterministic data point, not proof.