Bonsai 2 27B
RTX 5090 · PrismML llama.cpp fork · 262,144 ctx
- reported speed:
- 101.0 tokens/s generation
- quant:
- PQ2_0 (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
coding
User benchmarks Ternary Bonsai 2 27B at ~101 t/s on an RTX 5090 32GB, generating 106,396 tokens in ~17.5 minutes. Setup is a PrismML llama.cpp fork (release prism-b10683-d8f26ee, CUDA 12.8) with the PQ2_0 GGUF at 7.2 GB, 262,144 context, flash attention on, and reasoning effort xhigh. User compares it against Qwen 3.5 9B Q6_K at ~168 t/s and Gemma 4 12B Q8_0 at ~96 t/s on the same pagoda prompt, and rates Bonsai's output substantially better.