llamaperf

Qwen3.8 27B

on Intel Arc Pro B60 24GB · llama.cpp · 150,000 ctx

Tone: positive
Sep 26, 2026
Throughput
22.2 t/s gen · 245.5 t/s pp
Quant
Q4_K (GGUF)
VRAM reported
24 GB

Use cases

codingmathlong-context

Summary

User benchmarks a Qwen3.8-27B Ridge Intel Arc tuned GGUF on an Intel Arc Pro B60 24GB, reporting 22.23 tok/s pure autoregressive decode and 245.54 tok/s prompt prefill. Setup is llama.cpp with Q4_K weights, -ngl 99 and -fa on, on a single 24GB card. With embedded MTP speculative decoding the tuned Ridge build reaches 41.28 tok/s at 93.4% acceptance, and the DAS Lab IQ3_S tuned build reaches 41.88 tok/s at 92.9% acceptance. Stock IQ3_XXS decodes at 8.10 tok/s and stock Unsloth UD-Q4_K_S at 12.97 tok/s. At ~150k-token depth decode falls to 7.91 tok/s with DFlash2 and 7.21 tok/s with native MTP.