Qwen3.8 27B
on Intel Arc B580 12GB · FreeToken · 128 ctx
Sep 29, 2026
Summary
User reports Qwen3.8 27B UD-Q2_K_XL at 0.65 tok/s decode on an Intel Arc B580 12GB.
Setup is the FreeToken-Intel fork with PyTorch XPU plus native SYCL packed matvecs, context 128 and KV cache 256 tokens; TTFT 37.58 s and model load 47.39 s, with 9.47 GB model state.
The run used a fixed 16-token forced decode length and is presented as a development baseline, not a general performance claim; the same table lists Llama 3.2 1B Q4_K_M at 18.29 tok/s, Gemma 4 12B Q4_K_XL at 2.85 tok/s and Ternary-Bonsai 27B PQ2_0 at 2.10 tok/s.