llamaperf

BitNet b1.58

1 report

BitNet b1.58 VRAM requirements by size and quant →

How does BitNet b1.58 run on your hardware?

Pick your GPU or Mac, then read the speeds people reported for BitNet b1.58 on it, or estimate memory fit and speed.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Run BitNet b1.58 yourself? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance for BitNet b1.58

Filter this model’s reports by setup →
Thin page (1 of 3 reports needed for indexing). Add yours.

BitNet b1.58 2B

Unknown GPU · Project Zero

Tone: positive
reported speed:
36.0 tokens/s generation
quant:
ternary (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports BitNet b1.58-2B-4T at 36 t/s on a Xeon with Project Zero, a self-written C99 inference engine, about 1.8x faster than bitnet.cpp. The engine also runs SmolLM2 F16 at ~100 t/s on an i5-11300H, where llama.cpp is about 7% faster, and DeepSeek Q4_K at 1.9 t/s versus 13.7 t/s for llama.cpp. It is built with GCC and make, uses AVX-512 kernels for ternary packing, and exposes an OpenAI-compatible API with SSE streaming.

Oct 8, 2026

Get a weekly email of new BitNet b1.58 reports on any GPU.

Email me new reports