llamaperf

RTX 5070

NVIDIA · 12GB · 2 reports

See what fits on this GPU →

Use the calculator to check which models fit in 12 GB of VRAM. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →
This page is thin (2 of 3 reports needed for indexing). Help fill it in.
Tone: mixed
reported speed:
3.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User is considering Intel Arc B60 or B65 for Qwen 3.8 27B, currently getting 3 t/s on RTX 5070. Mentions a custom vLLM fork for Intel Arc achieving ~20 t/s on the larger B70.

reported speed:
22.0 tokens/s generation · 25.0 tokens/s prompt processing
quant:
IQ1_S (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports running Qwen3.8-Flash-Next with IQ1_S quant on a single RTX 5070 (12GB VRAM) using llama.cpp. Prompt processing speed 24.96 t/s, generation speed 22.0 t/s. Context length set to 10000.