Qwen3.8 27B
on Intel Arc Pro B60 24GB · llama.cpp · 512 ctx
Oct 3, 2026
Summary
User benchmarks Qwen3.8-27B at 15.34 t/s generation and 513.39 t/s prompt processing on a single Intel Arc Pro B60.
Setup is llama.cpp build b10452 with Vulkan backend, Q4_K_M GGUF weights and Q8_0 KV cache, full GPU offload with Flash Attention, at 512 tokens of context.
A second run on two Arc Pro B60 cards gives 14.99 t/s decode and 507.93 t/s prefill at 512 tokens; dual-card prefill scales to 643.16 t/s at 1k, 718.37 t/s at 2k and 149.30 t/s at 64k, while decode stays near 15 t/s.