Qwen3.8 27B Swift
on 2× Intel Arc Pro B60 24GB · llama.cpp · 131,072 ctx
Oct 3, 2026
Summary
User reports Qwen3.8-27B at 24.00 tok/s decode on dual Intel Arc Pro B60 24GB (48GB total).
Setup is llama.cpp build b11100 with SYCL F16 JIT, Q4_K_M GGUF, Q8_0 KV cache, 131072 context, and native embedded Q8_0 MTP draft speculative decoding.
Prefill is 609.47 tok/s at 512 tokens, scaling to 925.39 tok/s at 4k and 570.24 tok/s at 128k. Stock non-MTP decode is 16.13 tok/s; live llama-server chat with MTP ranges 20.00-30.29 tok/s.