llamaperf
Oct 3, 2026
Throughput
8.3 t/s gen · 112.6 t/s pp
Quant
Q4_K_XL (GGUF)

Summary

User benchmarks Qwen3.5 35B-A3B at 8.34 t/s generation and 112.62 t/s prompt processing on an Intel Arc Pro B60 24GB. Setup is llama.cpp build 8175 with the SYCL backend, Q4_K_XL GGUF weights, full GPU offload (-ngl 100). User reports the card is ok for chat-length context on smaller models but slow for agentic tools like opencode/claude code, where initial prompt response can take 5-10 minutes. A comparison run of the same model on an RTX Pro 4500 Blackwell reached 133.47 t/s generation and 3807.62 t/s prompt processing.