llamaperf

Unknown family

on AMD RX 6600 XT · llama.cpp · 130,416 ctx

Sep 26, 2026
Throughput
89.7 t/s pp
VRAM reported
8 GB

Summary

User reports prefill speed on an RX 6600 XT 8GB that does not fall steadily with context, using llama.cpp Vulkan with every layer offloaded: 935 t/s at about 1K tokens down to 318 t/s at 8K, then 472 t/s at 16K, and 89.7 t/s at 130K against 76.7 t/s at 65K. The model is not named. User says the spread across runs was under 2%.