llamaperf

Local model performance on your hardware

Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.

What runs on your hardware?

Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Local model performance reports from the community

These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →

GPU: Intel Arc A770 16GB
Compare setup details

Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.

Tone: positive
reported speed:
49.0 tokens/s generation · 48.0 tokens/s prompt processing
quant:
Q6_K (GGUF)
kv:
q4_1
flash attention:
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingtool-usevisionlong-contextagentic

User reports Qwen3.5-9B at 49 t/s generation and 48 t/s prompt processing on an Intel Arc A770 16GB. Setup is llama.cpp b9521 (Vulkan) with Q6_K weights, q4_1 KV cache, 256K context, flash attention, vision mmproj and MTP speculative decoding on a single slot. The same model under WSL2 with Q8_0 KV cache and 128K context reached about 40 t/s without vision. Qwopus3.5-4B-Coder hit 64 t/s generation and 100 t/s prompt at 96K context, and Gemma 4 12B reached 22 t/s generation and 74 t/s prompt at 128K context.

Oct 7, 2026
reported speed:
14.4 tokens/s generation
quant:
Q4_K_M (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen3.6-27B-A3B-Coder at 14.4 t/s decode on an Intel Arc A770 16GB, scoring 10/10 on a Lua CSV parser acceptance task. Setup is llama.cpp with the SYCL backend and GGUF Q4_K_M weights, a comparison point against arcint's own AWQ IR serving on the same card. The figure is described as a rough bound rather than a directly comparable measurement, since the quantisation and engine differ from arcint's production configuration.

Oct 5, 2026
reported speed:
33.2 tokens/s generation
quant:
Q8_0 (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks Qwen3-1.7B at 33.2 tok/s on an Intel Arc A770 16GB using llama.cpp with the SYCL backend. Setup is llama.cpp SYCL with Q8_0 quantisation, single-stream (n_parallel=1). OpenVINO via OVMS reached 65.4 tok/s on the same model. Ten models were tested in total, with OpenVINO faster than SYCL on every model.

Sep 23, 2026
Tone: mixed
reported speed:
12.0 tokens/s generation
quant:
Q3_K_S (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 27B at 12 t/s on an Intel Arc A770 using pure K-quants. Setup uses the SYCL backend, where I-quant mixes like Unsloth's UD-Q3_K_XL run at 7 t/s while pure K-quants like Bartowski's Q3_K_S run at 12 t/s. The user notes SYCL is bottlenecked by inefficient I-quant handling. The user is looking for the smallest K-quant-only Qwen3.8 27B at q3, with Bartowski's Q3_K_S at 12.7 GB as the current top contender.

Sep 14, 2026
Showing 1–4 of 4
Page 1 of 1

Community benchmarks snapshot

Records by GPU

NVIDIA RTX 3090165NVIDIA RTX 5090114AMD Strix Halo 128GB82NVIDIA DGX Spark57NVIDIA RTX 5060 Ti 16GB53NVIDIA RTX Pro 6000 Blackwell51NVIDIA RTX 3060 12GB46AMD Radeon AI PRO R9700 32GB43NVIDIA RTX 409036NVIDIA RTX 5070 Ti30

Records by model

1393 total
Qwen3.8778
Qwen3.6170
DeepSeek V4 Flash121
Gemma 461
Qwen3.529
Qwen322
other212

Records by engine

1070 total
llama.cpp569
vLLM153
Strata46
NInfer39
Ollama34
other229

Use cases

coding 450agentic 285long-context 207tool-use 120vision 85summarization 45math 36creative-writing 30multilingual 19text-generation 9rp 6reasoning 3
coding450agentic285long-context207tool-use120vision85summarization45math36creative-writing30

Median t/s by GPU

On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking

RTX 509093RTX Pro 600067M5 Ultra 256GB50RX 7900 XTX41RTX 309036V100 32GB33RTX 409032RX 7800 XT 16GB30RTX 5090 Laptop 24GB30Radeon AI PRO R9700 32GB29

Reports by model size

Qwen3.8 27B467Qwen3.8 125B · 6B active281DeepSeek V4 Flash 284B · 13B active102Qwen3.6 35B · 3B active100Qwen3.6 27B68Gemma 4 26B · 4B active25DeepSeek V4.1 Flash 552B · 16B active22Muse 30B18

Quants

Q4_K_M119NVFP495IQ4_XS57Q4_K_XL56UD-Q4_K_XL49IQ3_XXS38Q437Q8_0354-bit24MXFP423