llamaperf

Bonsai 2

PrismML · 1 report

Bonsai 2 VRAM requirements by size and quant →
Thin page (1 of 3 reports needed for indexing). Add yours.

Bonsai 2 27B

RTX 5090 · PrismML llama.cpp fork · 262,144 ctx

Tone: positive
reported speed:
101.0 tokens/s generation
quant:
PQ2_0 (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User benchmarks Ternary Bonsai 2 27B at ~101 t/s on an RTX 5090 32GB, generating 106,396 tokens in ~17.5 minutes. Setup is a PrismML llama.cpp fork (release prism-b10683-d8f26ee, CUDA 12.8) with the PQ2_0 GGUF at 7.2 GB, 262,144 context, flash attention on, and reasoning effort xhigh. User compares it against Qwen 3.5 9B Q6_K at ~168 t/s and Gemma 4 12B Q8_0 at ~96 t/s on the same pagoda prompt, and rates Bonsai's output substantially better.

Sep 18, 2026