Bonsai 2 27B
NVIDIA RTX 5060 Laptop 8GB · llama.cpp
- reported speed:
- 29.3 tokens/s generation
- quant:
- PTQ1_0 (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Ternary Bonsai 2 27B at 29.27 t/s on an RTX 5060 Laptop 8 GB under Windows, using a Prism llama.cpp build. Setup is llama.cpp with PTQ1_0 (5.53 GiB), temp 0, seed 42, a 2048 reasoning budget and a 3072 token cap on 50 MATH-500 questions. A −2 logit bias on "wait", "maybe" and "perhaps" (nine token ids) dropped accuracy from 44/50 to 43/50 and raised average tokens from 845.2 to 872.3; the biased run measured 29.28 t/s. The user calls it one deterministic data point, not proof.