Local model performance on your hardware
Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.
What runs on your hardware?
Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.
Free to use. No account needed. Memory estimates and community measurements are labelled separately.
Looking for a particular model?
Search for a model to see its reported speeds across GPUs and Macs.
Local model performance reports from the community
These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →
Compare setup details
Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.
No NVIDIA H100 NVL reports match these filters yet.
Estimate what fits on the NVIDIA H100 NVL →Share your own numbers →
Community benchmarks snapshot
Records by GPU
Records by model
Records by engine
Use cases
Median generation t/s by GPU
On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking