llamaperf

VRAM Calculator

Pick your GPU and what you use models for. Every open-weight model the site knows is sized for your card, graded on fit, given a speed estimate from the card's memory bandwidth, and ranked by how well it matches the job. Where community benchmarks exist on the same card, the measured number replaces the estimate.

16 GB unified memory·154 GB/s

Balanced: quality first, then speed.

Unified memory: the 16 GB pool is shared with the GPU, so there is no separate RAM to offload into.

6 perfect·17 good·1 marginal·19 too tight
Sort

Memory assumes an F16 KV cache at 32k context; Perfect is under 60% of the pool, Good under 85%, Marginal under 98%. Offload paths cap at Good. Speed is the card's memory bandwidth divided by the bytes read per token, at 55% efficiency, as llmfit estimates it; a community median at the same quant replaces it. Click a row for the breakdown.