Bonsai 2 27B
NVIDIA RTX 3080 10GB · llama.cpp · 32,768 ctx
- reported speed:
- 52.2 tokens/s generation · 1210.6 tokens/s prompt processing
- quant:
- PTQ1_0 (GGUF)
- kv:
- q8_0
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Bonsai 2 27B PTQ1_0 at 52.18 t/s generation and 1210.6 t/s prefill on an RTX 3080 10GB. Setup is llama.cpp (fork build prism-b10735-842b188, CUDA 12.8) with PTQ1_0 quant and q8_0 KV cache at 32K context, single card, batch 1, depth 0, -ngl 99 -fa on. User also benchmarks PQ2_0 at 61.38 t/s and Qwen3.8-27B-UD-IQ2_XXS at 44.46 t/s on the same card, and reports a context ladder up to 160K q4_0 at 9627 MiB peak VRAM. A LiveCodeBench v6 comparison gives PTQ1_0 32/50 pass@1 versus 16/50 for IQ2_XXS and 28/50 for PQ2_0.