llamaperf

M5 16GB

APPLE · 16GB unified memory · 1 report

See what fits on this GPU →

Use the calculator to check which models fit in 16 GB of unified memory. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →
This page is thin (1 of 3 reports needed for indexing). Help fill it in.

Qwen3.8 27B

M5 16GB · llama.cpp · 8,192 ctx

reported speed:
9.0 tokens/s generation
quant:
Q3_xxs
kv:
8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

summarization

User is new to local LLMs and asks if ~9 t/s is normal for this setup. They also ask for model recommendations for PDF summaries on 16GB RAM.