llamaperf

Qwen3.8 27B

on Intel Arc Pro B60 24GB · llama.cpp · 512 ctx

Oct 3, 2026
Throughput
15.3 t/s gen · 513.4 t/s pp
Quant
Q4_K_M (GGUF)
KV cache
Q8_0

Summary

User benchmarks Qwen3.8-27B at 15.34 t/s generation and 513.39 t/s prompt processing on a single Intel Arc Pro B60. Setup is llama.cpp build b10452 with Vulkan backend, Q4_K_M GGUF weights and Q8_0 KV cache, full GPU offload with Flash Attention, at 512 tokens of context. A second run on two Arc Pro B60 cards gives 14.99 t/s decode and 507.93 t/s prefill at 512 tokens; dual-card prefill scales to 643.16 t/s at 1k, 718.37 t/s at 2k and 149.30 t/s at 64k, while decode stays near 15 t/s.