Gemma 2 2B
on Unknown GPU · PULSAR-ASM
Oct 4, 2026
Summary
User reports Gemma-2B at 4.5-4.7 tokens/s in FP16 on an older quad-core i5 desktop, CPU only.
Setup is a custom 5.2 KB x86-64 assembly engine (PULSAR-ASM) using AVX2 and F16C with a 4-thread SMP GEMM for prefill, sustaining about 18.5 GB/s memory bandwidth on DDR4-2400.
The engine has zero C/C++ runtime and zero PyTorch dependencies; the Python harness only uses ctypes for VirtualAlloc and OS threads.