Qwen3.5 35B (3B active)
on Intel Arc Pro B60 24GB · llama.cpp
Oct 3, 2026
Summary
User benchmarks Qwen3.5 35B-A3B at 8.34 t/s generation and 112.62 t/s prompt processing on an Intel Arc Pro B60 24GB.
Setup is llama.cpp build 8175 with the SYCL backend, Q4_K_XL GGUF weights, full GPU offload (-ngl 100).
User reports the card is ok for chat-length context on smaller models but slow for agentic tools like opencode/claude code, where initial prompt response can take 5-10 minutes. A comparison run of the same model on an RTX Pro 4500 Blackwell reached 133.47 t/s generation and 3807.62 t/s prompt processing.