llamaperf

Gemma 2

2 reports

Gemma 2 VRAM requirements by size and quant →

How does Gemma 2 run on your hardware?

Pick your GPU or Mac, then read the speeds people reported for Gemma 2 on it, or estimate memory fit and speed.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Run Gemma 2 yourself? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance for Gemma 2

Filter this model’s reports by setup →
Thin page (2 of 3 reports needed for indexing). Add yours.

Gemma 2 2B

Unknown GPU · PULSAR-ASM

reported speed:
4.5-4.7 tokens/s generation
quant:
FP16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Gemma-2B at 4.5-4.7 tokens/s in FP16 on an older quad-core i5 desktop, CPU only. Setup is a custom 5.2 KB x86-64 assembly engine (PULSAR-ASM) using AVX2 and F16C with a 4-thread SMP GEMM for prefill, sustaining about 18.5 GB/s memory bandwidth on DDR4-2400. The engine has zero C/C++ runtime and zero PyTorch dependencies; the Python harness only uses ctypes for VirtualAlloc and OS threads.

Oct 4, 2026

Gemma 2 9B

Unknown GPU

Tone: positive
reported speed:
8.2 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Gemma 9B at 8.2 t/s sustained on a Snapdragon 8 Gen 5 phone, running purely on the CPU with an ~8GB memory footprint. Setup is a custom engine with ARM-specific optimizations, chosen over mobile OpenCL/Vulkan drivers, plus an ONNX-based Kokoro TTS voice engine running natively on-device. The run is fully offline and air-gapped with zero API calls.

Oct 3, 2026

Get a weekly email of new Gemma 2 reports on any GPU.

Email me new reports