llamaperf
Sep 23, 2026
Throughput
33.2 t/s gen
Quant
Q8_0 (GGUF)
System RAM
64 GB
VRAM reported
16 GB

Summary

User benchmarks Qwen3-1.7B at 33.2 tok/s on an Intel Arc A770 16GB using llama.cpp with the SYCL backend. Setup is llama.cpp SYCL with Q8_0 quantisation, single-stream (n_parallel=1). OpenVINO via OVMS reached 65.4 tok/s on the same model. Ten models were tested in total, with OpenVINO faster than SYCL on every model.