llamaperf

Qwen3.8 27B

on Intel Arc B580 12GB · FreeToken · 128 ctx

Sep 29, 2026
Throughput
0.7 t/s gen
Quant
UD-Q2_K_XL (GGUF)
VRAM reported
12 GB

Summary

User reports Qwen3.8 27B UD-Q2_K_XL at 0.65 tok/s decode on an Intel Arc B580 12GB. Setup is the FreeToken-Intel fork with PyTorch XPU plus native SYCL packed matvecs, context 128 and KV cache 256 tokens; TTFT 37.58 s and model load 47.39 s, with 9.47 GB model state. The run used a fixed 16-token forced decode length and is presented as a development baseline, not a general performance claim; the same table lists Llama 3.2 1B Q4_K_M at 18.29 tok/s, Gemma 4 12B Q4_K_XL at 2.85 tok/s and Ternary-Bonsai 27B PQ2_0 at 2.10 tok/s.