Qwen3 1.7B
on Intel Arc A770 16GB · llama.cpp
Sep 23, 2026
Summary
User benchmarks Qwen3-1.7B at 33.2 tok/s on an Intel Arc A770 16GB using llama.cpp with the SYCL backend.
Setup is llama.cpp SYCL with Q8_0 quantisation, single-stream (n_parallel=1).
OpenVINO via OVMS reached 65.4 tok/s on the same model. Ten models were tested in total, with OpenVINO faster than SYCL on every model.