llamaperf

Qwen3.8 27B Swift

on 2× Intel Arc Pro B60 24GB · llama.cpp · 131,072 ctx

Tone: positive
Oct 3, 2026
Throughput
24.0 t/s gen · 609.5 t/s pp
Quant
Q4_K_M (GGUF)
KV cache
Q8_0
VRAM reported
48 GB

Summary

User reports Qwen3.8-27B at 24.00 tok/s decode on dual Intel Arc Pro B60 24GB (48GB total). Setup is llama.cpp build b11100 with SYCL F16 JIT, Q4_K_M GGUF, Q8_0 KV cache, 131072 context, and native embedded Q8_0 MTP draft speculative decoding. Prefill is 609.47 tok/s at 512 tokens, scaling to 925.39 tok/s at 4k and 570.24 tok/s at 128k. Stock non-MTP decode is 16.13 tok/s; live llama-server chat with MTP ranges 20.00-30.29 tok/s.