Local model performance on your hardware
Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.
What runs on your hardware?
Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.
Free to use. No account needed. Memory estimates and community measurements are labelled separately.
Local model performance reports from the community
These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →
Model: DeepSeek V2
Compare setup details
Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.
No reports match these filters.
Community benchmarks snapshot
Records by GPU
Records by model
Qwen3.8793
Qwen3.6174
DeepSeek V4 Flash127
Gemma 463
Qwen3.534
Qwen325
other222
Records by engine
llama.cpp581
vLLM161
Strata48
NInfer41
Ollama35
other242
Use cases
coding 467agentic 296long-context 219tool-use 125vision 89summarization 47math 37creative-writing 31multilingual 19text-generation 9rp 6reasoning 3