llamaperf

Gemma 2 2B

on Unknown GPU · PULSAR-ASM

Oct 4, 2026
Throughput
4.5-4.7 t/s gen
Quant
FP16

Summary

User reports Gemma-2B at 4.5-4.7 tokens/s in FP16 on an older quad-core i5 desktop, CPU only. Setup is a custom 5.2 KB x86-64 assembly engine (PULSAR-ASM) using AVX2 and F16C with a 4-thread SMP GEMM for prefill, sustaining about 18.5 GB/s memory bandwidth on DDR4-2400. The engine has zero C/C++ runtime and zero PyTorch dependencies; the Python harness only uses ctypes for VirtualAlloc and OS threads.