llamaperf

M4 Max 64GB vs M6 24GB for local LLMs

7 reports on the M4 Max 64GB and 1 on the M6 24GB

Which is faster for local LLMs?

On Qwen3.8 27B at 4-bit, the M4 Max 64GB has 21.3 generation tokens/s (one run); prompt processing 143.7 tokens/s (one run); the M6 24GB has 9.8 generation tokens/s (one run). Those are community reports whose engines and settings differ, so the gap is a rough sign, not a controlled measure of the two cards. Memory bandwidth, which caps how fast a card can write, is 3.2x higher on the M4 Max 64GB (546 GB/s on the M4 Max 64GB, 171 on the M6 24GB).

M4 Max 64GB
Memory
64GB unified
Memory bandwidth
546 GB/s
FP16 compute
18.4 TFLOPS
Reports
7
M6 24GB
Memory
24GB unified
Memory bandwidth
171 GB/s
FP16 compute
not on record
Reports
1

Measured on both

Generation and prompt processing speeds from plain runs (one device, one request, no speculative decoding, the whole model in the device's memory) of the same model size at the same bit class. Each side shows the engines and reported context behind its median. Context can mean a configured window or prompt depth, so these are community comparisons with differing settings. How to read these

ModelM4 Max 64GBM6 24GB
Qwen3.8 27B4-bit21.3one runPP 143.7 (one run)engine not statedcontext not stated9.8one runmlx-vlmcontext not stated

Estimated by the calculator

The quant the calculator recommends for each card and its estimated speed at a 32,768-token context with 32 GB of system RAM, from memory bandwidth and the model's shape. Where a model was also measured above, the measurement wins.

ModelM4 Max 64GBM6 24GB
Qwen3.8 27B13.5 t/sQ8_016.8 t/sQ2_K
Qwen3.8 125B · 6B active61.4 t/sQ2_Kdoesn't fit at these settings
DeepSeek V4 Flash 284B · 13B activedoesn't fit at these settingsdoesn't fit at these settings
Qwen3.6 35B · 3B active42.5 t/sQ8_055.5 t/sQ2_K
Qwen3.6 27B13.5 t/sQ8_016.8 t/sQ2_K
Gemma 4 26B · 4B active36.4 t/sQ8_037.4 t/sQ3_K_S
GLM-5.3 320B · 18B activedoesn't fit at these settingsdoesn't fit at these settings

Change the context, RAM or model in the calculator: M4 Max 64GB · M6 24GB

Every report on each card, with its source: M4 Max 64GB · M6 24GB

More comparisons