llamaperf

Qwen2

1 report

Qwen2 VRAM requirements by size and quant →

How does Qwen2 run on your hardware?

Pick your GPU or Mac, then read the speeds people reported for Qwen2 on it, or estimate memory fit and speed.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Run Qwen2 yourself? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance for Qwen2

Filter this model’s reports by setup →
Thin page (1 of 3 reports needed for indexing). Add yours.

Qwen2 1.5B Function-Calling

Hailo-10H · hailo-llm-server · 2,048 ctx

reported speed:
7-10 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

tool-use

User reports Qwen2-1.5B function-calling model running on a Hailo-10H NPU with a Raspberry Pi 5, with decode speed of ~7–10 tok/s and time to first token of ~0.4–0.8s. Setup uses the hailo-llm-server engine via the hailo_platform.genai SDK, with a maximum context of 2048 tokens baked into the HEF. The server is single-tenant, serving one concurrent request; additional requests receive 429. The decode speed is given as a range.

Oct 4, 2026

Get a weekly email of new Qwen2 reports on any GPU.

Email me new reports