llamaperf

BitNet b1.58 2B

on Unknown GPU · Project Zero

Tone: positive
Oct 8, 2026
Throughput
36.0 t/s gen
Quant
ternary (GGUF)

Summary

User reports BitNet b1.58-2B-4T at 36 t/s on a Xeon with Project Zero, a self-written C99 inference engine, about 1.8x faster than bitnet.cpp. The engine also runs SmolLM2 F16 at ~100 t/s on an i5-11300H, where llama.cpp is about 7% faster, and DeepSeek Q4_K at 1.9 t/s versus 13.7 t/s for llama.cpp. It is built with GCC and make, uses AVX-512 kernels for ternary packing, and exposes an OpenAI-compatible API with SSE streaming.