Gemma 4 12B
AMD RX 6700 XT · llama.cpp · 8,192 ctx
- reported speed:
- 34.6 tokens/s generation · 653.9 tokens/s prompt processing
- quant:
- IQ4_NL (GGUF)
- kv:
- q8_0
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User benchmarks Gemma 4 12B (IQ4_NL, 6.24 GiB) on an AMD RX 6700 XT under llama.cpp, comparing the ROCm and Vulkan backends at 8192-token prefill and 512-token generation with q8_0 KV cache and flash-attention on. ROCm averages 653.9 t/s prefill and 34.60 t/s decode over 3 runs; Vulkan averages 354.4 t/s prefill and 40.92 t/s decode. ROCm is 84.5% faster on prefill but 15.4% slower on decode, giving a net wall-clock win of about 23% for a full 8192-prefill plus 512-generate cycle, with a crossover near 1760 prompt tokens. ROCm required two workarounds on gfx1031: building for gfx1030 with HSA_OVERRIDE_GFX_VERSION=10.3.0, and patching a flash-attention assert in fattn-common.cuh. A separate TOP_K sampler gap on ROCm is noted as under investigation.