llamaperf

VRAM Calculator

Pick your GPU or Mac and what you use models for. Every open-weight model the site knows is sized for your hardware, graded on fit, given a speed estimate from its memory bandwidth, and ranked by how well it matches the job. A Mac is graded against the memory macOS lets its GPU use, not all of it. Where community benchmarks exist on the same card, the measured number replaces the estimate.

The AMD Radeon 780M has no memory of its own and borrows the machine's RAM, so the calculator can't size models for it. Showing the NVIDIA RTX 4090 instead. Pick your hardware below, or see what people report on the AMD Radeon 780M.

Hardware

24 GB VRAM·32 GB system RAM·1,008 GB/s·330 TFLOPS fp16

Balanced: quality first, then speed.

Get a weekly email of new NVIDIA RTX 4090 reports and newly released models that fit it.

Email me new reports
25 perfect·41 good·1 marginal·25 too tight
Sort

Nobody has reported Gemma 4 12B on the NVIDIA RTX 4090 yet.

Memory assumes an F16 KV cache at 32k context; Perfect is under 60% of the pool, Good under 85%, Marginal under 98%. Offload paths cap at Good. Speed is the memory bandwidth divided by the bytes read per token, at 55% efficiency (llmfit's generic figure); a community median at the same quant replaces it. Click a row for the breakdown.

Did this help you choose a setup?

Optional feedback. No account needed.