llamaperf

RTX 4070

NVIDIA · 12GB · 4 reports

See what fits on this GPU →

Use the calculator to check which models fit in 12 GB of VRAM. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →

Qwen3.8 Flash-Next

RTX 4070 · Nebula · 24,576 ctx

reported speed:
7.2 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8-Flash-Next at 7.19 t/s on an RTX 4070 Ti 12GB with 128GB DDR4 RAM. Setup is the Nebula C/CUDA engine with native MTP speculative decoding, GPU expert caching (27 of 512 experts per layer resident in VRAM), and CPU MoE execution on an Intel i9-9940X using 14 threads, with a 24,576-token context capacity. A limited-tolerance acceptance mode reaches 7.54 t/s. Time to first token is 25.11 seconds for a 512-token input and 113.94 seconds for 2,048 tokens. Prefill is noted as slow for longer prompts.

Sep 12, 2026

Qwen3.8 27B

RTX 4070 · llama.cpp · 32,768 ctx

reported speed:
11.7 tokens/s generation
quant:
IQ4_XS (gguf)
kv:
Q8
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports 12.55 t/s for a 256 token response and 11.65 t/s for a 512 token response on an RPC split across an RTX 4070 Ti and an M5 MacBook Air with 16 GB unified memory. MTP off gives 9.72 t/s. The use case is agentic coding.

Sep 7, 2026

Qwen3.8 27B Uncensored

RTX 4070 · llama.cpp · 100,100 ctx

Tone: mixed
quant:
IQ4_XS (gguf)
kv:
Q4
rpcreative-writingvisionlong-context

User reports running Qwen3.8-27B on 16 GB of VRAM with the IQ4_XS quant. Setup is MTP disabled, mmproj offloaded to CPU, and a Q4 KV cache, at roughly 100k context. User notes Windows VRAM overhead as a constraint.

Aug 28, 2026
Tone: positive
reported speed:
55.0 tokens/s generation
quant:
Q4_K_M (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

text-generation

User reports about 55 t/s on an RTX 4070 12GB. The figure is estimated from compute-market tiers rather than a measured run.

May 1, 2026