llamaperf

Intel Arc Pro B70

INTEL · 12GB · 3 reports

See what fits on this GPU →

Use the calculator to check which models fit in 12 GB of VRAM. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →

Qwen3.8 27B

Intel Arc Pro B70 · vLLM · 128,000 ctx

Tone: positive
reported speed:
52.2 tokens/s generation · 763.0 tokens/s prompt processing
quant:
INT4 (GPTQ)
kv:
FP8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagenticvisiontool-use

Intel Arc Pro B70 32GB, vLLM XPU, Qwen3.8-27B GPTQ INT4, MTP2 speculative decoding, FP8 KV cache, 128K context. Median decode 52.2 tok/s, prefill 763 tok/s at 111.8K tokens. Vision, tool calling, and agent test pass. vLLM beats llama.cpp SYCL by ~1.8x. Caveats: MTP+concurrency crash fixed with max-num-seqs 1; prefix caching bug patched.

Qwen3.6 35B (3B active)

Intel Arc Pro B70 · llama.cpp · 262,000 ctx

Tone: positive
reported speed:
70.5 tokens/s generation · 977.4 tokens/s prompt processing
quant:
Q4_K_M (gguf)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports creating a poker game without issues. Also mentions trying Intel's vLLM fork previously.

reported speed:
63.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.