Gemma 4 12B
on NVIDIA RTX 5060 Laptop 8GB · llama.cpp · 4,096 ctx
Sep 24, 2026
Summary
User reports Gemma 4 12B at 40.0 t/s on an RTX 5060 Laptop GPU with 8 GB VRAM, using MTP speculative decoding with a Q6_K draft head and F16 KV cache at 4096 context.
Setup is llama.cpp with mainline draft-mtp support (commit b9193 or later), an imatrix-guided per-tensor fit-to-VRAM quant at 6.18 GB, 48 layers, parallel 1. Prefill is 700 to 900 t/s.
Without MTP the same setup reaches 27.4 t/s. With a Q8 KV cache it reaches 40.4 t/s. In real chat use the user sees 42 or more t/s on fresh context, dropping to 33 to 37 t/s with 16k of context filled.