Unknown family
AMD RX 6600 XT · llama.cpp · 130,416 ctx
- reported speed:
- 89.7 tokens/s prompt processing
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports prefill speed on an RX 6600 XT 8GB that does not fall steadily with context, using llama.cpp Vulkan with every layer offloaded: 935 t/s at about 1K tokens down to 318 t/s at 8K, then 472 t/s at 16K, and 89.7 t/s at 130K against 76.7 t/s at 65K. The model is not named. User says the spread across runs was under 2%.