llamaperf

Qwen2 1.5B Function-Calling

on Hailo-10H · hailo-llm-server · 2,048 ctx

Oct 4, 2026
Throughput
7-10 t/s gen

Use cases

tool-use

Summary

User reports Qwen2-1.5B function-calling model running on a Hailo-10H NPU with a Raspberry Pi 5, with decode speed of ~7–10 tok/s and time to first token of ~0.4–0.8s. Setup uses the hailo-llm-server engine via the hailo_platform.genai SDK, with a maximum context of 2048 tokens baked into the HEF. The server is single-tenant, serving one concurrent request; additional requests receive 429. The decode speed is given as a range.