Qwen3.8
on AMD Radeon 780M · llama.cpp
Sep 7, 2026
Summary
User reports ROCm 7.14 on a Radeon 780M with llama.cpp is fast but unstable and crashes.
Setting AMD_SERIALIZE_KERNEL=3 stabilizes it but drops prompt processing to about 100 t/s, down from 200-300 t/s without the workaround.
User asks for advice.