Unknown family
on AMD RX 6600 XT · llama.cpp · 130,416 ctx
Sep 26, 2026
Summary
User reports prefill speed on an RX 6600 XT 8GB that does not fall steadily with context, using llama.cpp Vulkan with every layer offloaded: 935 t/s at about 1K tokens down to 318 t/s at 8K, then 472 t/s at 16K, and 89.7 t/s at 130K against 76.7 t/s at 65K.
The model is not named. User says the spread across runs was under 2%.