llamaperf

Gemma 4 12B

on AMD RX 6700 XT · llama.cpp · 8,192 ctx

Tone: mixed
Oct 6, 2026
Throughput
34.6 t/s gen · 653.9 t/s pp
Quant
IQ4_NL (GGUF)
KV cache
q8_0
System RAM
16 GB
VRAM reported
12 GB

Use cases

long-context

Summary

User benchmarks Gemma 4 12B (IQ4_NL, 6.24 GiB) on an AMD RX 6700 XT under llama.cpp, comparing the ROCm and Vulkan backends at 8192-token prefill and 512-token generation with q8_0 KV cache and flash-attention on. ROCm averages 653.9 t/s prefill and 34.60 t/s decode over 3 runs; Vulkan averages 354.4 t/s prefill and 40.92 t/s decode. ROCm is 84.5% faster on prefill but 15.4% slower on decode, giving a net wall-clock win of about 23% for a full 8192-prefill plus 512-generate cycle, with a crossover near 1760 prompt tokens. ROCm required two workarounds on gfx1031: building for gfx1030 with HSA_OVERRIDE_GFX_VERSION=10.3.0, and patching a flash-attention assert in fattn-common.cuh. A separate TOP_K sampler gap on ROCm is noted as under investigation.