llamaperf

M4 32GB

APPLE · 32GB unified memory · 2 reports

See what fits on this GPU →

Use the calculator to check which models fit in 32 GB of unified memory. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →All M4 Macs compared →
This page is thin (2 of 3 reports needed for indexing). Help fill it in.
Tone: positive
reported speed:
22.0 tokens/s generation
quant:
Q4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports 8-22 t/s on a 32GB M4 MacBook Air with 21GB of allocations. Setup is a custom inference engine called Cherenkov that combines predictive expert streaming with optional mixed-precision execution, keeping a bounded working set of experts in unified memory rather than loading the entire model. A one-layer lookahead predicts which experts will be needed next and initiates SSD reads.

Sep 11, 2026
Tone: positive
reported speed:
8.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User asks about an MTP or DFlash head for Qwen 3.8 27B, and mentions 35B-A3B as a daily driver. User reports about 8 tok/s for Qwen 3.6 27B with MTP on 32 GB of unified memory.

Sep 7, 2026