llamaperf

Qwen3.8

on AMD Radeon 780M · llama.cpp

Tone: mixed
Sep 7, 2026
Throughput
100.0 t/s pp

Summary

User reports ROCm 7.14 on a Radeon 780M with llama.cpp is fast but unstable and crashes. Setting AMD_SERIALIZE_KERNEL=3 stabilizes it but drops prompt processing to about 100 t/s, down from 200-300 t/s without the workaround. User asks for advice.