Local model performance on your hardware
Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.
What runs on your hardware?
Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.
Free to use. No account needed. Memory estimates and community measurements are labelled separately.
Local model performance reports from the community
These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →
Model: Qwen2
Compare setup details
Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.
No reports match these filters.
Community benchmarks snapshot
Records by GPU
Records by model
Qwen3.8781
Qwen3.6170
DeepSeek V4 Flash121
Gemma 461
Qwen3.529
Qwen322
other213
Records by engine
llama.cpp570
vLLM153
Strata47
NInfer40
Ollama34
other230
Use cases
coding 453agentic 287long-context 208tool-use 120vision 85summarization 45math 36creative-writing 30multilingual 19text-generation 9rp 6reasoning 3