llamaperf

Local model performance on your hardware

Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.

What runs on your hardware?

Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Local model performance reports from the community

These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →

GPU: M5 Ultra 256GB
Compare setup details (1 active)

Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.

reported speed:
76.7 tokens/s generation · 2270.0 tokens/s prompt processing
quant:
oQ4 (MLX)
kv:
4-bit

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks oMLX 0.7.0 against 0.7.0rc1 on an M5 Ultra 256GB, running GLM-5.3-Flash oQ4 at 4K–200K context. Prefill improves from 792 to 2,270 tok/s (median) and decode from 55.4 to 76.7 tok/s. A 1M-token prompt completes with prefill 1,485 tok/s, decode 41.5 tok/s, and time to first token 11.8 min. Setup uses Lightning MTP, TurboQuant KV 4-bit, and one model loaded at a time.

Oct 6, 2026
reported speed:
68.8 tokens/s generation
quant:
oQ4e

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports GLM-5.3 Flash at 68.8 tok/s on an M5 Ultra 256GB. Setup is oMLX 0.7.0 with oQ4e quantization and MTP speculative decoding, with more optimization still to come. Prefill is 1,878 toks. User shared the result because other benchmarks looked lower than expected.

Oct 5, 2026
reported speed:
73.4 tokens/s generation · 1010.0 tokens/s prompt processing
quant:
Q3-G128-LSQ (MLX)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User asks whether oMLX benchmarks can be spoofed, citing a third-party oMLX benchmark entry for deepseek-v4.1-flash-q3-g128-lsq-mlx on an M5 Ultra with 256GB RAM, posted two days earlier, reporting about 73.4 tok/s generation and 1,010 tok/s prompt processing. The run is not the user's own; the figures come from a linked oMLX benchmark page and screenshot.

Oct 3, 2026
Tone: positive
reported speed:
29.0 tokens/s generation
quant:
oQ8e (MLX)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codinglong-context

User reports Qwen3.8-Flash-Next 8-bit at 29 t/s decode on a 1M-token prompt on an M5 Ultra 256GB Mac Studio. Setup is oMLX 0.7.0 with oQ8e weights and YaRN x4 position scaling; the 1M-token first read took 5.3 minutes and the next turn 6.7 s, finding all three hidden codes. At 250k context the model prefills at 4,235 t/s and decodes at 63 t/s; the user notes YaRN stretches a 262k-trained model and that a needle test does not prove reasoning at 1M.

Oct 3, 2026
Tone: positive
reported speed:
59-74 tokens/s generation
quant:
oQ8e (MLX)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User compares Qwen3.8-Flash-Next against Laguna-S-2.1 on a Mac Studio M5 Ultra 256GB, both capped at 262K context with thinking on and unique content per run. Qwen runs under oMLX with an oQ8e quant and MTP speculation, holding roughly 4,200 tok/s prefill and 59-74 tok/s decode across sizes. Laguna runs under LM Studio with an 8-bit quant, prefill degrading superlinearly from 10.3s at 8K to 455.4s at 200K, and decode falling from 68 to 34 tok/s with no speculation. Quality was a draw at 4/4 each on four script-verified problems. At 200K, prefill is 94% of total time on both. The user notes an earlier run was invalidated by shared prefixes letting the KV cache carry over.

Sep 30, 2026
Tone: positive
reported speed:
108.0 tokens/s generation · 2887.0 tokens/s prompt processing
quant:
oQ4e (MLX)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agentic

User reports Qwen3.8-Flash-Next at 108 t/s generation and 2,887 t/s prompt processing on an M5 Ultra 256GB Mac Studio. Setup is oMLX 0.7.0.dev2 with MLX backend, oQ4e 4-bit quant, multi-token prediction depth 3, 16K prompt. M3 Ultra 512GB comparison: 70 t/s generation, 1,143 t/s prompt. At 256K context, M5 prompt processing was 2,544 t/s. Qwen3.8-27B on M5 Ultra: 48 t/s generation, 1,701 t/s prompt; RTX 5090 PC in LM Studio: 59 t/s generation, 3,031 t/s prompt.

Sep 29, 2026

Qwen3.8 27B

M5 Ultra 256GB · oMLX · 8,192 ctx

Tone: positive
reported speed:
50.0 tokens/s generation · 1800.0 tokens/s prompt processing
quant:
q4
mtp (multi-token prediction):
off

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 27B at 50 t/s generation and 1800 t/s prompt processing on an M5 Ultra at 8k context. Figures come from the omlx website rather than a local run, with q4 weights and no MTP. The user notes the benchmarks' provenance is unclear but finds them reasonable. User calls the results very promising.

Sep 17, 2026
Showing 1–7 of 7
Page 1 of 1

Community benchmarks snapshot

Records by GPU

NVIDIA RTX 3090165NVIDIA RTX 5090114AMD Strix Halo 128GB82NVIDIA DGX Spark57NVIDIA RTX 5060 Ti 16GB53NVIDIA RTX Pro 6000 Blackwell51NVIDIA RTX 3060 12GB46AMD Radeon AI PRO R9700 32GB43NVIDIA RTX 409037NVIDIA RTX 5070 Ti30

Records by model

1397 total
Qwen3.8781
Qwen3.6170
DeepSeek V4 Flash121
Gemma 461
Qwen3.529
Qwen322
other213

Records by engine

1074 total
llama.cpp570
vLLM153
Strata47
NInfer40
Ollama34
other230

Use cases

coding 453agentic 287long-context 208tool-use 120vision 85summarization 45math 36creative-writing 30multilingual 19text-generation 9rp 6reasoning 3
coding453agentic287long-context208tool-use120vision85summarization45math36creative-writing30

Median t/s by GPU

On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking

RTX 509093RTX Pro 600067M5 Ultra 256GB50RX 7900 XTX41RTX 309036V100 32GB33RTX 409032RX 7800 XT 16GB30RTX 5090 Laptop 24GB30Radeon AI PRO R9700 32GB29

Reports by model size

Qwen3.8 27B468Qwen3.8 125B · 6B active283DeepSeek V4 Flash 284B · 13B active102Qwen3.6 35B · 3B active100Qwen3.6 27B68Gemma 4 26B · 4B active25DeepSeek V4.1 Flash 552B · 16B active22GLM-5.3 320B · 18B active18

Quants

Q4_K_M119NVFP495IQ4_XS57Q4_K_XL56UD-Q4_K_XL49IQ3_XXS38Q437Q8_0354-bit24Q6_K23