llamaperf

NVIDIA RTX 4070 Ti

NVIDIA · 12GB · 1 report

As of 8 Oct 2026, the models most run on the NVIDIA RTX 4070 Ti, with the median of plain runs (one device, one request, no speculative decoding):

Engines people use on it: llama.cpp 1

Run models on your NVIDIA RTX 4070 Ti? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the NVIDIA RTX 4070 Ti

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 12 GB of VRAM.

This page is thin (1 of 3 reports needed for indexing). Help fill it in.

Qwen3.8 27B

NVIDIA RTX 4070 Ti · llama.cpp · 4,096 ctx

reported speed:
43.6 tokens/s generation · 943.3 tokens/s prompt processing
quant:
UD-IQ2_XXS (GGUF)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks Qwen3.8-27B at 43.643 generation tok/s and 943.278 prompt tok/s on an RTX 4070 Ti 12GB. Setup is llama.cpp b10448 with UD-IQ2_XXS GGUF, Q8 KV cache, 4096-token context, all 66 layers on GPU, one slot, 256-token outputs. A second quant UD-Q2_K_XL decoded 38.030 tok/s and 995.095 prompt tok/s with 1,583 MiB more peak VRAM. An IQ2 context ladder gave 41.124 tok/s at 4K, 40.522 at 8K and 39.201 at 16K (12,831 prompt tokens, 11.119 s TTFT). An isolated IQ2 draft-mtp experiment raised throughput 47.284% on prose and 92.651% on Python code, but prose output diverged at token 16, so MTP stays off by default. In a 24-task quality pass, Q2 scored 10/24 and IQ2 9/24 (McNemar p = 1.0).

Sep 28, 2026

Get a weekly email of new NVIDIA RTX 4070 Ti reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Qwen3.8 27B
NVIDIA RTX 4070 Ti
UD-IQ2_XXS
llama.cpp
4,09643.6 tokens/s