Bonsai 2 27B
on NVIDIA RTX 3080 10GB · llama.cpp · 32,768 ctx
Sep 28, 2026
Use cases
coding
Summary
User reports Bonsai 2 27B PTQ1_0 at 52.18 t/s generation and 1210.6 t/s prefill on an RTX 3080 10GB.
Setup is llama.cpp (fork build prism-b10735-842b188, CUDA 12.8) with PTQ1_0 quant and q8_0 KV cache at 32K context, single card, batch 1, depth 0, -ngl 99 -fa on.
User also benchmarks PQ2_0 at 61.38 t/s and Qwen3.8-27B-UD-IQ2_XXS at 44.46 t/s on the same card, and reports a context ladder up to 160K q4_0 at 9627 MiB peak VRAM. A LiveCodeBench v6 comparison gives PTQ1_0 32/50 pass@1 versus 16/50 for IQ2_XXS and 28/50 for PQ2_0.