llamaperf

NVIDIA RTX 4070 Super

NVIDIA · 12GB · 2 reports

Engines people use on it: ik_llama.cpp 1 · llama.cpp 1

Run models on your NVIDIA RTX 4070 Super? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the NVIDIA RTX 4070 Super

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 12 GB of VRAM.

This page is thin (2 of 3 reports needed for indexing). Help fill it in.

Qwen3.8 27B

NVIDIA RTX 4070 Super · llama.cpp · 64,000 ctx

Tone: positive
reported speed:
18.0 tokens/s generation · 500-600 tokens/s prompt processing
quant:
IQ3_S (GGUF)
kv:
q4_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8-27B at ~18 tok/s decode and ~500-600 tok/s prefill at 64k context on an RTX 4070 Super 12GB with 32GB system RAM. Setup is stock llama.cpp with an ISTA-DASLab GSQ-RCO IQ3_S quant (~11GB), q4_0 KV cache, -ngl 58 and token_embd offloaded to CPU. Only 16 of 64 layers need KV cache, so 64k context takes about 1.1GB instead of 4GB at f16. The same recipe works with 0bserverx' Qwen3.8-27B-Heretic-GSQ-RCO IQ3_S quant at about a 7% speed loss. User calls GSQ-RCO the best Qwen3.8-27B quant tested and says it performs close to full precision.

Oct 7, 2026

Qwen3.6 35B (3B active)

NVIDIA RTX 4070 Super · ik_llama.cpp · 131,072 ctx

Tone: positive
reported speed:
110.2 tokens/s generation
quant:
IQ4_XS-4.19bpw (GGUF)
kv:
Q8
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingsummarizationmath

User benchmarks llama.cpp at 89.76 t/s against ik_llama.cpp at 110.24 t/s on Qwen3.6-35B-A3B with MTP. Setup is the IQ4_XS quant on a Ryzen 7 9700X running CachyOS, with the GPU used as a secondary card and the iGPU for display. The ik_llama.cpp run is a 23% speed increase.

May 21, 2026

Get a weekly email of new NVIDIA RTX 4070 Super reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Qwen3.8 27B
NVIDIA RTX 4070 Super
IQ3_S
llama.cpp
64,00018.0 tokens/s
Qwen3.6 35B (3B active)
NVIDIA RTX 4070 Super
IQ4_XS-4.19bpw
ik_llama.cpp
131,072110.2 tokens/s