llamaperf

RTX 3060 Laptop 6GB

NVIDIA · 6GB · 1 report

See what fits on this GPU →

Use the calculator to check which models fit in 6 GB of VRAM. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →
This page is thin (1 of 3 reports needed for indexing). Help fill it in.

Bonsai 27B

RTX 3060 Laptop 6GB · mentria.ai · 3,072 ctx

Tone: positive
reported speed:
30.0 tokens/s generation
quant:
1-bit

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports 25-30 t/s decode in a chat UI and 32 t/s raw in a WebGPU browser inference engine, with prompt processing of 1489 tokens in about 25 seconds. The model is natively 1-bit, 27B parameters in 3.8 GB, running at 3072 tokens of context on a 6 GB card. The user also mentions smaller tiers (Qwen3.5 0.8B, 2B, 4B) and a vision tower.

Sep 9, 2026