llamaperf
Oct 5, 2026
Throughput
14.4 t/s gen
Quant
Q4_K_M (GGUF)
VRAM reported
16 GB

Use cases

coding

Summary

User reports Qwen3.6-27B-A3B-Coder at 14.4 t/s decode on an Intel Arc A770 16GB, scoring 10/10 on a Lua CSV parser acceptance task. Setup is llama.cpp with the SYCL backend and GGUF Q4_K_M weights, a comparison point against arcint's own AWQ IR serving on the same card. The figure is described as a rough bound rather than a directly comparable measurement, since the quantisation and engine differ from arcint's production configuration.