llamaperf

Bonsai-27B

1 report

Thin page (1 of 3 reports needed for indexing). Add yours.

Bonsai 27B

RTX 3060 Laptop 6GB · mentria.ai · 3,072 ctx

Tone: positive
reported speed:
30.0 tokens/s generation
quant:
1-bit

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Browser inference engine in WebGPU. Decode 25-30 tok/s in chat UI, raw 32 tok/s. Prompt processing 1489 tokens in ~25s. Context 3072 tokens on 6GB card. Model is natively 1-bit, 27B params in 3.8GB. Also mentions smaller tiers (Qwen3.5 0.8B, 2B, 4B) and vision tower.