llamaperf

Intel Arc A770 16GB

INTEL · 16GB · 4 reports

As of 7 Oct 2026, the models most run on the Intel Arc A770 16GB, with the median of plain runs (one device, one request, no speculative decoding):

Engines people use on it: llama.cpp 3

Run models on your Intel Arc A770 16GB? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the Intel Arc A770 16GB

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 16 GB of VRAM.

Tone: positive
reported speed:
49.0 tokens/s generation · 48.0 tokens/s prompt processing
quant:
Q6_K (GGUF)
kv:
q4_1
flash attention:
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingtool-usevisionlong-contextagentic

User reports Qwen3.5-9B at 49 t/s generation and 48 t/s prompt processing on an Intel Arc A770 16GB. Setup is llama.cpp b9521 (Vulkan) with Q6_K weights, q4_1 KV cache, 256K context, flash attention, vision mmproj and MTP speculative decoding on a single slot. The same model under WSL2 with Q8_0 KV cache and 128K context reached about 40 t/s without vision. Qwopus3.5-4B-Coder hit 64 t/s generation and 100 t/s prompt at 96K context, and Gemma 4 12B reached 22 t/s generation and 74 t/s prompt at 128K context.

Oct 7, 2026
reported speed:
14.4 tokens/s generation
quant:
Q4_K_M (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen3.6-27B-A3B-Coder at 14.4 t/s decode on an Intel Arc A770 16GB, scoring 10/10 on a Lua CSV parser acceptance task. Setup is llama.cpp with the SYCL backend and GGUF Q4_K_M weights, a comparison point against arcint's own AWQ IR serving on the same card. The figure is described as a rough bound rather than a directly comparable measurement, since the quantisation and engine differ from arcint's production configuration.

Oct 5, 2026
reported speed:
33.2 tokens/s generation
quant:
Q8_0 (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks Qwen3-1.7B at 33.2 tok/s on an Intel Arc A770 16GB using llama.cpp with the SYCL backend. Setup is llama.cpp SYCL with Q8_0 quantisation, single-stream (n_parallel=1). OpenVINO via OVMS reached 65.4 tok/s on the same model. Ten models were tested in total, with OpenVINO faster than SYCL on every model.

Sep 23, 2026
Tone: mixed
reported speed:
12.0 tokens/s generation
quant:
Q3_K_S (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 27B at 12 t/s on an Intel Arc A770 using pure K-quants. Setup uses the SYCL backend, where I-quant mixes like Unsloth's UD-Q3_K_XL run at 7 t/s while pure K-quants like Bartowski's Q3_K_S run at 12 t/s. The user notes SYCL is bottlenecked by inefficient I-quant handling. The user is looking for the smallest K-quant-only Qwen3.8 27B at q3, with Bartowski's Q3_K_S at 12.7 GB as the current top contender.

Sep 14, 2026

Get a weekly email of new Intel Arc A770 16GB reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Qwen3.5 9B DeepSeek-V4-Flash
Intel Arc A770 16GB
Q6_K
llama.cpp
262,14449.0 tokens/s
Qwen3.6 27B (3B active) Coder
Intel Arc A770 16GB
Q4_K_M
llama.cpp
Not reported14.4 tokens/s
Qwen3 1.7B
Intel Arc A770 16GB
Q8_0
llama.cpp
Not reported33.2 tokens/s
Qwen3.8 27B
Intel Arc A770 16GB
Q3_K_S
Engine not reported
Not reported12.0 tokens/s