llamaperf

Gemma 4 12B

on NVIDIA RTX 5060 Laptop 8GB · llama.cpp · 4,096 ctx

Sep 24, 2026
Throughput
40.0 t/s gen · 700-900 t/s pp
Quant
Q6_K (GGUF)
KV cache
F16
VRAM reported
8 GB

Summary

User reports Gemma 4 12B at 40.0 t/s on an RTX 5060 Laptop GPU with 8 GB VRAM, using MTP speculative decoding with a Q6_K draft head and F16 KV cache at 4096 context. Setup is llama.cpp with mainline draft-mtp support (commit b9193 or later), an imatrix-guided per-tensor fit-to-VRAM quant at 6.18 GB, 48 layers, parallel 1. Prefill is 700 to 900 t/s. Without MTP the same setup reaches 27.4 t/s. With a Q8 KV cache it reaches 40.4 t/s. In real chat use the user sees 42 or more t/s on fresh context, dropping to 33 to 37 t/s with 16k of context filled.