llamaperf

RTX 5070 Ti Laptop 12GB

NVIDIA · 12GB · 3 reports

See what fits on this GPU →

Use the calculator to check which models fit in 12 GB of VRAM. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →

Qwen3.8 27B

RTX 5070 Ti Laptop 12GB · Unsloth Studio · 8,192 ctx

Tone: positive
reported speed:
4.5 tokens/s generation · 23.5 tokens/s prompt processing
quant:
UD-Q4_K_XL
kv:
Q8
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

MTP acceptance 78-83%. Longer response 3.26 t/s, short factual 4.42 t/s, coding 4.53 t/s. Prompt processing 20-27 t/s. Model size ~17.9GB, offloading to CPU.

Tone: positive
reported speed:
59.0 tokens/s generation · 155.8 tokens/s prompt processing
quant:
Q4_K_M (gguf)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingmath

Compared Qwen3.8 27B dense at Q2 and Q3 against Qwen3.6 35B-A3B MoE on 12GB VRAM. MoE was fastest and passed sanity test; dense Q3 was slow (7.5-9.1 t/s). Author prefers MoE for local use.

Qwen3.8 27B

RTX 5070 Ti Laptop 12GB · llama.cpp · 262,144 ctx

Tone: positive
reported speed:
5.0 tokens/s generation
quant:
Q4_K_XL (gguf)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User runs Qwen 3.8 27B UD Q4_K_XL on RTX 5070 Ti Mobile 12GB with llama.cpp. Two configs: context-prioritized (ctx 262144, ~1.5-5 t/s) and speed-prioritized (ctx 98304, ~9-11.5 t/s). KV cache q8_0. Praises model's capability and instruction following.