llamaperf

AMD MI50 16GB

AMD · 16GB · 1 report

See what fits on this GPU →

Use the calculator to check which models fit in 16 GB of VRAM. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →
This page is thin (1 of 3 reports needed for indexing). Help fill it in.
Tone: positive
reported speed:
16.9 tokens/s generation
quant:
IQ4 (GGUF)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 Flash-Next at 16.9 t/s on a single MI50 16GB with the gfx906-16gb-expert-pool llama.cpp fork. Setup uses the moe-expert-pool branch with IQ4 GGUF weights, 128K context and Q8 KV cache; the expert cache profile of 66 admits 144 pools at roughly 69% hit rate and peaks at 15.24 GiB VRAM. A stock CPU MoE path without the expert cache reached 11.76 t/s, and up to 19.8 t/s was observed once warm.

Sep 18, 2026