Thin page (2 of 3 reports needed for indexing).
Add yours.
- reported speed:
- 9.1 tokens/s generation
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Bonsai 1.7B at about 9.1 t/s on an Intel N97 CPU drawing 12 watts.
The model solved simple physics problems in this run.
- reported speed:
- 30.0 tokens/s generation
- quant:
- 1-bit
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports 25-30 t/s decode in a chat UI and 32 t/s raw in a WebGPU browser inference engine, with prompt processing of 1489 tokens in about 25 seconds.
The model is natively 1-bit, 27B parameters in 3.8 GB, running at 3072 tokens of context on a 6 GB card.
The user also mentions smaller tiers (Qwen3.5 0.8B, 2B, 4B) and a vision tower.