llamaperf

M4 16GB

APPLE · 16GB unified memory · 1 report

See what fits on this GPU →

Use the calculator to check which models fit in 16 GB of unified memory. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →All M4 Macs compared →
This page is thin (1 of 3 reports needed for indexing). Help fill it in.

Maple Preview 20B (1B active)

M4 16GB · Mference · 131,072 ctx

Tone: positive
reported speed:
20.0 tokens/s generation · 40.0 tokens/s prompt processing
quant:
ternary

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Maple Preview, a 20B A1B MoE model trained from scratch in ternary precision, running on a MacBook Air M4 with 16 GB RAM. Setup is Mference, a fork of turbo-fieldfare, streaming experts from SSD and reducing memory usage to 500-1200 MB. The model has limited world knowledge and multilingual capabilities but can use tools. The user is enthusiastic about the low memory footprint and potential for background agentic workloads.

Sep 7, 2026