Local model performance on your hardware
Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.
What runs on your hardware?
Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.
Free to use. No account needed. Memory estimates and community measurements are labelled separately.
Local model performance reports from the community
These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →
Model: Qwen3-VL
Compare setup details
Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.
No reports match these filters.
Community benchmarks snapshot
Records by GPU
Records by model
Qwen3.8820
Qwen3.6177
DeepSeek V4 Flash128
Gemma 466
Qwen3.537
Qwen326
other234
Records by engine
llama.cpp593
vLLM166
Strata52
NInfer41
Ollama36
other253
Use cases
coding 483agentic 311long-context 228tool-use 127vision 92summarization 47math 37creative-writing 31multilingual 19text-generation 9rp 6reasoning 3