Gemma 2 2B
Unknown GPU · PULSAR-ASM
- reported speed:
- 4.5-4.7 tokens/s generation
- quant:
- FP16
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Gemma-2B at 4.5-4.7 tokens/s in FP16 on an older quad-core i5 desktop, CPU only. Setup is a custom 5.2 KB x86-64 assembly engine (PULSAR-ASM) using AVX2 and F16C with a 4-thread SMP GEMM for prefill, sustaining about 18.5 GB/s memory bandwidth on DDR4-2400. The engine has zero C/C++ runtime and zero PyTorch dependencies; the Python harness only uses ctypes for VirtualAlloc and OS threads.