llamaperf

Intel Arc Pro B60 24GB

INTEL · 24GB · 5 reports

As of 7 Oct 2026, the models most run on the Intel Arc Pro B60 24GB, with the median of plain runs (one device, one request, no speculative decoding):

Engines people use on it: llama.cpp 5

Run models on your Intel Arc Pro B60 24GB? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the Intel Arc Pro B60 24GB

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 24 GB of VRAM.

Tone: positive
reported speed:
91.9 tokens/s generation · 1760.0 tokens/s prompt processing
quant:
Q4_0 (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticcoding

User reports Nemotron 3.5 Lightning 30B-A3B at 91.91 tok/s decode on 2x Intel Arc Pro B60 24GB. Setup is llama.cpp SYCL with Q4_0 weights and an MTP Q8_0 drafter at --spec-draft-n-max 7, 22.18 GiB VRAM, 1,760 tok/s prefill at 12K context. MTP acceptance is 99.5-100% at every n-max; the model needs the whole card with no co-residence.

Oct 3, 2026

Qwen3.8 27B

Intel Arc Pro B60 24GB · llama.cpp · 512 ctx

reported speed:
15.3 tokens/s generation · 513.4 tokens/s prompt processing
quant:
Q4_K_M (GGUF)
kv:
Q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks Qwen3.8-27B at 15.34 t/s generation and 513.39 t/s prompt processing on a single Intel Arc Pro B60. Setup is llama.cpp build b10452 with Vulkan backend, Q4_K_M GGUF weights and Q8_0 KV cache, full GPU offload with Flash Attention, at 512 tokens of context. A second run on two Arc Pro B60 cards gives 14.99 t/s decode and 507.93 t/s prefill at 512 tokens; dual-card prefill scales to 643.16 t/s at 1k, 718.37 t/s at 2k and 149.30 t/s at 64k, while decode stays near 15 t/s.

Oct 3, 2026

Qwen3.8 27B Swift

2× Intel Arc Pro B60 24GB · llama.cpp · 131,072 ctx

Tone: positive
reported speed:
24.0 tokens/s generation · 609.5 tokens/s prompt processing
quant:
Q4_K_M (GGUF)
kv:
Q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8-27B at 24.00 tok/s decode on dual Intel Arc Pro B60 24GB (48GB total). Setup is llama.cpp build b11100 with SYCL F16 JIT, Q4_K_M GGUF, Q8_0 KV cache, 131072 context, and native embedded Q8_0 MTP draft speculative decoding. Prefill is 609.47 tok/s at 512 tokens, scaling to 925.39 tok/s at 4k and 570.24 tok/s at 128k. Stock non-MTP decode is 16.13 tok/s; live llama-server chat with MTP ranges 20.00-30.29 tok/s.

Oct 3, 2026
Tone: mixed
reported speed:
8.3 tokens/s generation · 112.6 tokens/s prompt processing
quant:
Q4_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks Qwen3.5 35B-A3B at 8.34 t/s generation and 112.62 t/s prompt processing on an Intel Arc Pro B60 24GB. Setup is llama.cpp build 8175 with the SYCL backend, Q4_K_XL GGUF weights, full GPU offload (-ngl 100). User reports the card is ok for chat-length context on smaller models but slow for agentic tools like opencode/claude code, where initial prompt response can take 5-10 minutes. A comparison run of the same model on an RTX Pro 4500 Blackwell reached 133.47 t/s generation and 3807.62 t/s prompt processing.

Oct 3, 2026

Qwen3.8 27B

Intel Arc Pro B60 24GB · llama.cpp · 150,000 ctx

Tone: positive
reported speed:
22.2 tokens/s generation · 245.5 tokens/s prompt processing
quant:
Q4_K (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingmathlong-context

User benchmarks a Qwen3.8-27B Ridge Intel Arc tuned GGUF on an Intel Arc Pro B60 24GB, reporting 22.23 tok/s pure autoregressive decode and 245.54 tok/s prompt prefill. Setup is llama.cpp with Q4_K weights, -ngl 99 and -fa on, on a single 24GB card. With embedded MTP speculative decoding the tuned Ridge build reaches 41.28 tok/s at 93.4% acceptance, and the DAS Lab IQ3_S tuned build reaches 41.88 tok/s at 92.9% acceptance. Stock IQ3_XXS decodes at 8.10 tok/s and stock Unsloth UD-Q4_K_S at 12.97 tok/s. At ~150k-token depth decode falls to 7.91 tok/s with DFlash2 and 7.21 tok/s with native MTP.

Sep 26, 2026

Get a weekly email of new Intel Arc Pro B60 24GB reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Nemotron 3.5 Lightning 30B (3B active)
2× Intel Arc Pro B60 24GB
Q4_0
llama.cpp
Not reported91.9 tokens/s
Qwen3.8 27B
Intel Arc Pro B60 24GB
Q4_K_M
llama.cpp
51215.3 tokens/s
Qwen3.8 27B Swift
2× Intel Arc Pro B60 24GB
Q4_K_M
llama.cpp
131,07224.0 tokens/s
Qwen3.5 35B (3B active)
Intel Arc Pro B60 24GB
Q4_K_XL
llama.cpp
Not reported8.3 tokens/s
Qwen3.8 27B
Intel Arc Pro B60 24GB
Q4_K
llama.cpp
150,00022.2 tokens/s