llamaperf

Local model performance on your hardware

Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.

What runs on your hardware?

Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Local model performance reports from the community

These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →

GPU: Intel Arc Pro B60 24GB
Compare setup details

Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.

Tone: positive
reported speed:
91.9 tokens/s generation · 1760.0 tokens/s prompt processing
quant:
Q4_0 (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticcoding

User reports Nemotron 3.5 Lightning 30B-A3B at 91.91 tok/s decode on 2x Intel Arc Pro B60 24GB. Setup is llama.cpp SYCL with Q4_0 weights and an MTP Q8_0 drafter at --spec-draft-n-max 7, 22.18 GiB VRAM, 1,760 tok/s prefill at 12K context. MTP acceptance is 99.5-100% at every n-max; the model needs the whole card with no co-residence.

Oct 3, 2026

Qwen3.8 27B

Intel Arc Pro B60 24GB · llama.cpp · 512 ctx

reported speed:
15.3 tokens/s generation · 513.4 tokens/s prompt processing
quant:
Q4_K_M (GGUF)
kv:
Q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks Qwen3.8-27B at 15.34 t/s generation and 513.39 t/s prompt processing on a single Intel Arc Pro B60. Setup is llama.cpp build b10452 with Vulkan backend, Q4_K_M GGUF weights and Q8_0 KV cache, full GPU offload with Flash Attention, at 512 tokens of context. A second run on two Arc Pro B60 cards gives 14.99 t/s decode and 507.93 t/s prefill at 512 tokens; dual-card prefill scales to 643.16 t/s at 1k, 718.37 t/s at 2k and 149.30 t/s at 64k, while decode stays near 15 t/s.

Oct 3, 2026

Qwen3.8 27B Swift

2× Intel Arc Pro B60 24GB · llama.cpp · 131,072 ctx

Tone: positive
reported speed:
24.0 tokens/s generation · 609.5 tokens/s prompt processing
quant:
Q4_K_M (GGUF)
kv:
Q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8-27B at 24.00 tok/s decode on dual Intel Arc Pro B60 24GB (48GB total). Setup is llama.cpp build b11100 with SYCL F16 JIT, Q4_K_M GGUF, Q8_0 KV cache, 131072 context, and native embedded Q8_0 MTP draft speculative decoding. Prefill is 609.47 tok/s at 512 tokens, scaling to 925.39 tok/s at 4k and 570.24 tok/s at 128k. Stock non-MTP decode is 16.13 tok/s; live llama-server chat with MTP ranges 20.00-30.29 tok/s.

Oct 3, 2026
Tone: mixed
reported speed:
8.3 tokens/s generation · 112.6 tokens/s prompt processing
quant:
Q4_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks Qwen3.5 35B-A3B at 8.34 t/s generation and 112.62 t/s prompt processing on an Intel Arc Pro B60 24GB. Setup is llama.cpp build 8175 with the SYCL backend, Q4_K_XL GGUF weights, full GPU offload (-ngl 100). User reports the card is ok for chat-length context on smaller models but slow for agentic tools like opencode/claude code, where initial prompt response can take 5-10 minutes. A comparison run of the same model on an RTX Pro 4500 Blackwell reached 133.47 t/s generation and 3807.62 t/s prompt processing.

Oct 3, 2026

Qwen3.8 27B

Intel Arc Pro B60 24GB · llama.cpp · 150,000 ctx

Tone: positive
reported speed:
22.2 tokens/s generation · 245.5 tokens/s prompt processing
quant:
Q4_K (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingmathlong-context

User benchmarks a Qwen3.8-27B Ridge Intel Arc tuned GGUF on an Intel Arc Pro B60 24GB, reporting 22.23 tok/s pure autoregressive decode and 245.54 tok/s prompt prefill. Setup is llama.cpp with Q4_K weights, -ngl 99 and -fa on, on a single 24GB card. With embedded MTP speculative decoding the tuned Ridge build reaches 41.28 tok/s at 93.4% acceptance, and the DAS Lab IQ3_S tuned build reaches 41.88 tok/s at 92.9% acceptance. Stock IQ3_XXS decodes at 8.10 tok/s and stock Unsloth UD-Q4_K_S at 12.97 tok/s. At ~150k-token depth decode falls to 7.91 tok/s with DFlash2 and 7.21 tok/s with native MTP.

Sep 26, 2026
Showing 1–5 of 5
Page 1 of 1

Community benchmarks snapshot

Records by GPU

NVIDIA RTX 3090165NVIDIA RTX 5090114AMD Strix Halo 128GB82NVIDIA DGX Spark57NVIDIA RTX 5060 Ti 16GB53NVIDIA RTX Pro 6000 Blackwell51NVIDIA RTX 3060 12GB46AMD Radeon AI PRO R9700 32GB43NVIDIA RTX 409036NVIDIA RTX 5070 Ti30

Records by model

1393 total
Qwen3.8778
Qwen3.6170
DeepSeek V4 Flash121
Gemma 461
Qwen3.529
Qwen322
other212

Records by engine

1070 total
llama.cpp569
vLLM153
Strata46
NInfer39
Ollama34
other229

Use cases

coding 450agentic 285long-context 207tool-use 120vision 85summarization 45math 36creative-writing 30multilingual 19text-generation 9rp 6reasoning 3
coding450agentic285long-context207tool-use120vision85summarization45math36creative-writing30

Median t/s by GPU

On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking

RTX 509093RTX Pro 600067M5 Ultra 256GB50RX 7900 XTX41RTX 309036V100 32GB33RTX 409032RX 7800 XT 16GB30RTX 5090 Laptop 24GB30Radeon AI PRO R9700 32GB29

Reports by model size

Qwen3.8 27B467Qwen3.8 125B · 6B active281DeepSeek V4 Flash 284B · 13B active102Qwen3.6 35B · 3B active100Qwen3.6 27B68Gemma 4 26B · 4B active25DeepSeek V4.1 Flash 552B · 16B active22Muse 30B18

Quants

Q4_K_M119NVFP495IQ4_XS57Q4_K_XL56UD-Q4_K_XL49IQ3_XXS38Q437Q8_0354-bit24MXFP423